What is the difference between ByteBuffer.allocate() and ByteBuffer.allocateDirect()?
answer
- heap = byte[] on GC heap; direct = native off-heap
- direct skips the extra I/O copy (zero-copy)
- direct: costly to allocate, freed by Cleaner not GC
- rule: direct for large/long-lived/channel I/O
- direct.hasArray() == false
basics
~20 sallocate() makes a buffer backed by a normal Java byte array on the heap. allocateDirect() makes a buffer in off-heap native memory, which the OS can use for I/O without an extra copy, but it is slower to create.
solid answer
~40 sByteBuffer.allocate(n) returns a heap buffer wrapping a plain byte[] managed by the garbage collector. allocateDirect(n) returns a direct buffer whose storage lives outside the Java heap in native memory. The key tradeoff is I/O: when you read/write a channel, the OS works with native memory, so a heap buffer must be copied into a temporary direct buffer first; a direct buffer skips that copy (zero-copy). However, direct buffers are expensive to allocate and free (the memory is reclaimed via Cleaner/PhantomReference, not normal GC), so allocating many short-lived ones is wasteful. Use direct buffers for large, long-lived buffers doing heavy channel I/O; use heap buffers for everything else, especially small or short-lived ones. Direct buffers also can't expose a backing array via array().
code
java · 12 linesByteBuffer heap = ByteBuffer.allocate(1024); // byte[] on the Java heap
ByteBuffer direct = ByteBuffer.allocateDirect(1024); // native off-heap memory
System.out.println(heap.isDirect()); // false
System.out.println(heap.hasArray()); // true -> heap.array() works
System.out.println(direct.isDirect()); // true
System.out.println(direct.hasArray()); // false -> direct.array() throws
// Heavy, reused channel read buffer: prefer direct
try (var ch = java.nio.channels.FileChannel.open(java.nio.file.Path.of("data.bin"))) {
ch.read(direct); // OS reads straight into native memory, no extra copy
}go deeper
Knows allocate() is on-heap and allocateDirect() is off-heap, and that direct is for fast I/O but costly to create.
Explains the extra copy heap buffers incur on channel I/O, that direct buffers are zero-copy, and states the large/long-lived rule of thumb.
Adds the GC-relocation reason the OS needs native memory, Cleaner-based reclamation, native-memory exhaustion risk, and pooling direct buffers.
Reasons about -XX:MaxDirectMemorySize tuning, off-heap leak diagnosis, when zero-copy actually pays off vs. allocation churn, and pooling strategies at system scale.
## Background: what a ByteBuffer is Java NIO (`java.nio`, "New I/O", added in Java 1.4) provides `Buffer` classes for moving bytes to and from **channels** (files, sockets). A `ByteBuffer` is a fixed-size container of bytes with a position, limit, and capacity. There are two physical ways its bytes are stored, and that is exactly what this topic is about. ## The Java heap vs. native (off-heap) memory - The **Java heap** is the memory region the JVM manages and the **garbage collector (GC)** scans. A normal `byte[]` lives here. The GC can move objects around to compact the heap. - **Native memory** (sometimes "off-heap") is memory the JVM requests directly from the operating system, outside the GC-managed heap. The GC does not scan it and cannot move it. ## allocate() — a heap buffer ```java ByteBuffer heap = ByteBuffer.allocate(1024); ``` This creates a **heap buffer**: internally it is backed by a `byte[1024]` on the Java heap. `heap.hasArray()` returns `true` and `heap.array()` gives you that array. It is cheap to create (just a normal object allocation) and freed automatically by GC. The catch is **I/O**. The operating system's read/write system calls require a stable, native memory address. A heap array can be moved by the GC at any moment, so the JVM cannot safely hand its address to the OS. So when you do a channel read/write with a **heap** buffer, the JVM internally allocates a temporary **direct** buffer, copies the bytes through it, then copies into your heap buffer. That is an extra copy on every I/O call. ## allocateDirect() — a direct buffer ```java ByteBuffer direct = ByteBuffer.allocateDirect(1024); ``` This creates a **direct buffer** whose bytes live in **native memory** at a fixed address. Because the address is stable, the OS can read/write it directly — **no intermediate copy** ("zero-copy" relative to the heap-buffer path). `direct.isDirect()` is `true`; `direct.hasArray()` is `false` (calling `array()` throws). The costs: - **Allocation/free is expensive.** Getting native memory from the OS is far slower than a heap object, and the memory is reclaimed not by ordinary GC but by a `Cleaner` (a `PhantomReference`) that runs only after the buffer object becomes unreachable and GC notices. Until then the native memory stays held, so leaks and out-of-memory in native space are easy if you churn many direct buffers. - Not subject to the normal `-Xmx` heap limit; instead governed by `-XX:MaxDirectMemorySize`. ## The rule of thumb Use **direct** buffers for buffers that are **large and long-lived and do heavy channel I/O** (e.g. a reusable read buffer in a server). Use **heap** buffers for **small, short-lived, or non-I/O** work. Allocating thousands of tiny direct buffers is an anti-pattern; pool them instead. ## Related operations in this area - `ByteBuffer.wrap(existingArray)` — does **not** allocate; it makes a heap buffer that *reuses* an existing `byte[]`, so writes through the buffer mutate that array. - View buffers like `asIntBuffer()`, `asLongBuffer()` — reinterpret the same underlying ByteBuffer's bytes as a sequence of larger primitives (sharing storage, not copying). - `asReadOnlyBuffer()` — returns a view that throws on any put; useful to hand out data safely. - `ByteBuffer.order(ByteOrder)` — sets big-endian (the default, network order) vs little-endian for multi-byte reads/writes. Knowing both kinds, the copy that heap buffers incur, and the allocation cost of direct buffers, lets you answer: *direct when the buffer is large, reused, and channel-bound; heap otherwise.*
- Why can't the OS read directly into a heap buffer's array?The garbage collector may move (relocate) heap objects to compact memory, so a heap array has no stable native address. A blocking OS read could outlive a GC move and corrupt memory, so the JVM copies through a stable temporary direct buffer instead.
- How is a direct buffer's memory reclaimed?Not by ordinary GC of the bytes. A Cleaner / PhantomReference attached to the buffer object runs after the buffer becomes unreachable and GC processes it, then frees the native memory. This is non-deterministic, which is why churning direct buffers can exhaust native memory before they're cleaned.
A heap buffer is luggage inside the airport (the JVM); to load it on the plane (the OS) it must first be moved to a cart on the tarmac (a temp direct buffer). A direct buffer already sits on the tarmac — loaded directly, but it took longer to get there and you must remember to remove it.
saying these in an interview costs you the question
- Claiming direct buffers are always faster — they are slower to allocate/free and only win on heavy channel I/O.
- Thinking direct buffer memory counts against -Xmx (it's native, bounded by -XX:MaxDirectMemorySize).
- Allocating a new direct buffer per request/iteration instead of pooling/reusing.
- Believing direct.array() works — it throws UnsupportedOperationException.