Why is System.arraycopy typically faster than copying an array with a manual for-loop?
answer
- JIT intrinsic → bulk/SIMD memory move
- one range check vs per-element bounds checks
- biggest win on large primitive arrays
- memmove semantics → overlap-safe
- loop only wins for transforms or tiny copies
basics
~20 sSystem.arraycopy is implemented inside the JVM as a single block memory copy, so it moves many bytes at once instead of one element per loop iteration. A hand-written for-loop does an array bounds check and a separate store on every element.
solid answer
~50 sSystem.arraycopy is a native method that the JIT compiler treats as an intrinsic: instead of a Java loop, the JVM emits an optimized bulk memory move (often using vectorized/SIMD instructions or a single hardware memcpy) and skips per-element array bounds checks after one upfront range check. A manual for-loop, by contrast, does a bounds check and an individual load/store each iteration; the JIT can sometimes vectorize and hoist those checks, but arraycopy gives you that optimized path guaranteed and portably. The gap is largest for big primitive arrays. For very small copies (a handful of elements) the difference is negligible and the call overhead can even make a loop competitive. So for anything but trivial sizes, prefer arraycopy: it is faster, clearer, and correctly handles overlapping ranges when source and destination are the same array.
go deeper
Knows arraycopy is faster because the JVM does the copy in bulk rather than element by element.
Can explain per-element bounds checks vs a single range check and that arraycopy moves a block of memory.
Describes it as a JIT intrinsic emitting vectorized/memcpy code, knows the small-copy and transform exceptions, and the memmove overlap guarantee.
Reasons quantitatively about cache behavior, SIMD width, write barriers for reference arrays, and when to benchmark rather than assume; weighs intrinsic guarantees vs JIT loop optimizations.
## Setup: what "copying an array" costs To copy N elements with a plain loop you write: ```java for (int i = 0; i < n; i++) { dest[i] = src[i]; } ``` Naively, each iteration does several things: read `src[i]` (a **bounds check** that `i < src.length` plus a memory load), write `dest[i]` (another bounds check plus a store), increment `i`, and test the loop condition. That is per-element work repeated N times. ## What an *intrinsic* is The **JIT (Just-In-Time) compiler** turns hot bytecode into machine code at runtime. An **intrinsic** is a method the JIT recognizes by name and replaces with a hand-tuned machine-code template instead of compiling its actual body. `System.arraycopy` is one of the most important intrinsics in the JVM. ## Why the intrinsic wins 1. **One range check, not N.** The intrinsic validates the whole `[srcPos, srcPos+length)` and `[destPos, destPos+length)` ranges **once** up front, then copies without re-checking each element. The naive loop pays a bounds check per element (the JIT can sometimes eliminate these, but not always). 2. **Bulk / vectorized move.** The copy is done as a **block memory operation** — frequently mapped to a hardware `memmove`/`memcpy` or SIMD (Single Instruction, Multiple Data) instructions that move 16/32/64 bytes per instruction instead of one element at a time. 3. **Cache- and width-friendly.** The native code copies in word- or cache-line-sized chunks, using the memory bandwidth efficiently rather than issuing one narrow store per element. 4. **No reference-write barrier surprises.** For primitive arrays there are no GC card-marking write barriers; for reference arrays the intrinsic handles the necessary barriers in a batched, optimized way. ## When the loop is competitive - **Tiny copies** (a few elements): the fixed overhead of the intrinsic call/dispatch can equal or exceed a short unrolled loop, so the difference is in the noise. - **Element transformation**: if you need to *change* values while copying (e.g. `dest[i] = src[i] * 2`), arraycopy cannot help — it only moves bytes. Use a loop (or streams) then. ## Overlap correctness A subtle bonus: when `src == dest` and the ranges overlap, `System.arraycopy` behaves as if it copied through a temporary buffer (`memmove` semantics), so it produces the correct result even when shifting elements right. A naive forward loop would **corrupt** data on a right shift because it would overwrite source elements before reading them. ## Practical guidance For any non-trivial pure copy, use `System.arraycopy` (or `Arrays.copyOf`, which wraps it): it is faster on large primitive arrays, equally correct on small ones, clearer in intent, and overlap-safe. Reserve manual loops for when you must transform elements during the copy.
- Can a manual for-loop ever match System.arraycopy's speed?For small copies the difference is negligible, and a JIT can sometimes vectorize and hoist bounds checks out of a simple copy loop to approach it. But arraycopy gives the optimized intrinsic path guaranteed and portably, so it is the safer default for large copies.
- Why is System.arraycopy safe for overlapping ranges when source equals destination?It has memmove semantics — it behaves as if the source range were first copied to a temporary buffer, so overlapping shifts (e.g. shifting elements right to make room) produce correct results, unlike a naive forward-writing loop.
saying these in an interview costs you the question
- Claiming arraycopy is always faster — for tiny copies the difference is negligible
- Thinking arraycopy can transform elements during the copy (it only moves bytes)
- Saying a forward loop is always correct for in-place right shifts — it corrupts overlapping data
- Ignoring that the JIT, not the source, makes arraycopy fast (it is an intrinsic)