skip to content

A tracing garbage collector in the JVM is free to relocate objects, but an operating-system read or write call needs a memory address that stays fixed for the whole duration of the call. How does the JVM bridge that gap when Java code hands a heap-allocated byte[] or a heap-backed java.nio.ByteBuffer to native code or to an NIO channel — what are the pinning and copying options, and what does each one cost?

level: seniorimportance: should knowfreq 30%

answer

  1. Moving collector → no address to hand the kernel
  2. Only two bridges: pin or copy
  3. GetPrimitiveArrayCritical → region pinning or GCLocker stall
  4. NIO always copies via per-thread temporary direct buffer
  5. jdk.nio.maxCachedBufferSize caps the per-thread high-water mark

basics

~20 s

Relocating collectors move objects, so a heap array has no stable address a system call can use. The JVM either pins it — JNI critical sections, which stall relocation — or copies it. NIO always copies, through a per-thread cached native buffer.

solid answer

~50 s

A system call keeps touching the address you gave it while it runs; a compacting or evacuating collector may move the backing `byte[]` at any point. So the JVM never hands a raw heap-object address to the kernel. Two bridges exist. **Pin.** `GetPrimitiveArrayCritical` in JNI asks for the array's real memory. Depending on the collector, HotSpot either pins the containing region or blocks relocation outright via the **GCLocker**, deferring any GC until the last critical section exits (you see `GCLocker Initiated GC` in the log). Cheap for tiny sections; a long or blocking one stalls every allocating thread. **Copy.** `java.nio` always copies: a heap `ByteBuffer` passed to a channel is copied into a **temporary direct buffer** taken from a per-thread cache, the syscall runs against that stable native address, and reads copy back. Correct and pause-free, but it costs one copy per operation, and the cached buffer stays alive per thread sized to that thread's largest transfer — bound it with `jdk.nio.maxCachedBufferSize`.

code

text · 7 lines
text
# A collection that was requested while a JNI critical section held the GCLocker,
# and only ran once the last critical thread exited:
[12.418s][info][gc] GC(37) Pause Young (Concurrent Start) (GCLocker Initiated GC) 812M->190M(2048M) 41.2ms

# Bound the per-thread temporary direct buffer cache used for heap-buffer I/O.
# Buffers larger than this are freed after use instead of retained per thread:
-Djdk.nio.maxCachedBufferSize=262144

go deeper

for a junior

Know the core fact: Java objects can be moved by the garbage collector, so their addresses are not stable, and that is why byte arrays cannot be handed straight to the operating system.

for a middle

Name both bridges and pick correctly: pinning via JNI critical sections, or copying — and know that java.nio always copies heap buffers into native memory before a channel operation.

for a senior

Explain the costs as operational symptoms: GCLocker-deferred collections and allocation stalls with no GC in the log on the pinning side; per-operation copies plus a per-thread native high-water mark, bounded by jdk.nio.maxCachedBufferSize, on the copying side.

for a principal

Frame it as an interface-design decision: NIO chose an unconditional copy to keep the collector unconstrained, trading throughput for isolation, while JNI critical sections push the constraint onto callers who rarely honour it. Drive the architecture toward pooled native buffers so neither bridge is on the hot path, and treat the FFM critical linker option as the same bargain with the risk finally made explicit.

## The constraint: managed objects have no stable address A `byte[]` is an ordinary Java object. Every mainstream HotSpot collector — Serial, Parallel, G1, Shenandoah, ZGC — is a *moving* collector: it reclaims space by evacuating live objects into fresh regions and rewriting the references that point at them. G1 and Parallel do it during a stop-the-world pause; Shenandoah and ZGC do it *concurrently*, while application threads are running, using barriers to redirect accesses to the new copy. A system call such as `read(2)` or `write(2)` takes a raw address and a length, and the kernel may touch that memory at any instant until the call returns. Nothing about a system call participates in read barriers or reference rewriting. Handing it the current address of a movable object is therefore not merely slow — it is a use-after-move memory-corruption bug waiting for the next collection. This is the single constraint that shapes the whole heap-to-native boundary in the JVM. There are exactly two ways out: make the object stop moving (**pin**), or make a stable copy the kernel can own (**copy**). ## The pinning path: JNI critical sections and the GCLocker JNI's `GetByteArrayElements` is deliberately vague: it returns a pointer and an `isCopy` out-parameter, letting the implementation copy or pin as it pleases. `GetPrimitiveArrayCritical` / `ReleasePrimitiveArrayCritical` is the stronger request — "give me the actual bytes, do not copy" — and it comes with a contract the *caller* must honour: between the two calls the thread must not block and must not call arbitrary JNI functions. What HotSpot does inside that window is collector-dependent, and the JNI spec still permits an outright copy: - Collectors that support **region pinning** mark the region containing the array as un-evacuable and let collection proceed around it. G1 gained this, removing its dependence on the global mechanism below for critical sections. - Otherwise the runtime falls back to the **GCLocker**: while any thread is inside a critical section, garbage collections that would relocate objects are *blocked*. A GC requested during that window is recorded and triggered as soon as the last critical thread leaves, which is why gc logs show entries attributed to a GCLocker-initiated collection. The cost model follows directly. A one-microsecond critical section around a checksum is free. A critical section that spans an entire file write — or, worse, blocks on I/O — pins the collector for that whole duration. Every thread that needs to allocate and cannot get a collection simply waits. The failure looks bizarre from the outside: allocation stalls with plenty of free heap, and long pause-like latency spikes with no corresponding GC work. ## The copying path: NIO's temporary direct buffers `java.nio` refuses to play that game at all. Channel implementations only ever issue system calls against **native** memory. When you pass a heap-backed `ByteBuffer` to `channel.write(...)`, the implementation obtains a temporary direct buffer from a **per-thread cache**, copies your payload into it, performs the syscall against that fixed native address, and — for reads — copies the result back into your array. The cache keeps a small number of buffers per thread and satisfies a request with the smallest cached buffer large enough; if none fits, it allocates a new one and may release the largest. That design has two consequences people meet in production: 1. **A per-operation copy proportional to payload size.** For large or high-rate transfers this is real CPU and memory bandwidth, and it is precisely the cost that using an already-native buffer avoids. 2. **A per-thread native high-water mark.** The cache is keyed to the thread and sized to the *largest* transfer that thread has ever performed. One 64 MB write on each of 200 pool threads can retain 200 × 64 MB of native memory for the life of those threads. It is not a leak in the code's sense — it is a cache — but it presents exactly like one: RSS far above `-Xmx`, unaffected by heap dumps, nothing visible in the heap. The system property `jdk.nio.maxCachedBufferSize` caps the size of buffer eligible for caching; anything bigger is released after use rather than retained. Setting it is standard hygiene for services that do occasional large transfers on many threads. ## Where the native side actually lives The memory a direct buffer wraps is obtained from the process allocator, not from the Java heap. The Java object holds an address; the payload bytes are never scanned, never relocated, and are not counted against `-Xmx`. Reclamation is post-mortem — attached cleanup runs once the wrapper object becomes unreachable — rather than anything the tracing collector does to the bytes themselves. For the boundary question that matters for one reason only: the address is **stable by construction**, so no pin and no copy is required. ## The same tradeoff, restated by the FFM API The foreign function and memory API (final in JDK 22) makes the choice explicit rather than implicit. A `MemorySegment` may be *native* or *heap*-backed. Downcalls normally require native segments for pointer arguments; passing a heap segment is allowed only through the linker's `critical(allowHeapAccess)` option, which is documented as suitable for short, non-blocking native calls — the same critical-section bargain as JNI, with the danger written into the API surface instead of into a spec footnote. ## Practical rules Keep critical sections tiny and non-blocking, or do not use them. If a buffer is used for I/O repeatedly, allocate it natively once and reuse it, so neither bridge is needed. If you must pass heap buffers, accept the copy and bound the per-thread cache. And when native footprint grows with thread count and transfer size while the heap is calm, suspect the temporary direct-buffer cache before suspecting a leak.

  • A service shows resident memory far above -Xmx, heap dumps look healthy, and the growth tracks the size of its thread pool. What is your first hypothesis and how do you test it?
    The per-thread temporary direct-buffer cache that NIO uses when heap-backed ByteBuffers are written to channels: each thread retains a native buffer sized to its largest-ever transfer. Test it by enabling Native Memory Tracking and by correlating footprint with the largest single transfer size times the number of I/O threads. Confirm by setting jdk.nio.maxCachedBufferSize to a small bound and watching the footprint flatten, then fix properly by using pooled direct buffers or chunking large transfers.
  • Why does the JNI spec allow GetPrimitiveArrayCritical to copy anyway, if the whole point is to avoid a copy?
    Because the JNI contract is written against many possible collectors, and not every collector can cheaply make a region un-evacuable. The spec guarantees only that the returned pointer is valid for the duration of the critical section, leaving the implementation free to pin, to make the region non-moving, or to hand back a copy. That is why correct native code must not assume writes through the pointer are visible before ReleasePrimitiveArrayCritical, and must always call the release even on error paths.
  • If you already own the buffer, what removes both the pin and the copy from the picture entirely?
    Allocating the payload in native memory in the first place — a direct buffer or a native MemorySegment — so its address is fixed by construction and the kernel can work against it in place. Because such allocation is expensive, high-throughput networking libraries pool and reuse these buffers rather than allocating per request, which keeps the native footprint bounded and predictable as well.

The heap is a warehouse whose staff reshuffle shelves while you work. A courier needs a fixed loading bay. You can either freeze all reshuffling while the courier walks in (pinning — nobody else gets served meanwhile), or carry the parcel out to the bay yourself first (copying — an extra trip every time, and the bay space you claimed stays reserved for you).

saying these in an interview costs you the question

  • Saying the JVM just passes the byte[] address to the syscall and the GC "knows not to move it"
  • Treating a JNI critical section as free — not realising it can defer collections for every thread in the process
  • Assuming GetPrimitiveArrayCritical is guaranteed never to copy
  • Explaining growing native memory in an NIO service as a direct-buffer leak, when it is the per-thread temporary buffer cache doing exactly what it was designed to do
  • Believing the copy for heap buffers exists for safety of the Java code rather than because the collector relocates objects

context