A thread's private allocation buffer in the Java heap no longer has room for the next object. What choices does the HotSpot JVM make at that point, and what is meant by allocation-buffer waste?
answer
- slow path: retire+refill vs allocate outside
- waste = abandoned tail of retired buffer
- TLABWasteTargetPercent default ~1%
- filler int[] preserves heap parsability
- large objects never fit a buffer
basics
~20 sIt either retires the buffer and grabs a fresh one, or allocates that one object in the shared young space. Retiring abandons the unused tail, which is waste; HotSpot only retires when the remaining tail is below a threshold, otherwise the big object goes outside the buffer.
solid answer
~60 sWhen the bump fails the bounds check, the thread enters the slow path and weighs two options. If the space left in the buffer is small — below a refill-waste limit derived from `TLABWasteTargetPercent`, about 1% of buffer size by default — the buffer is retired: the leftover tail is filled with a dummy object so the heap stays walkable, and the thread takes a fresh buffer from the shared young space using an atomic operation. That abandoned tail is the waste. If the remaining space is large, throwing it away would be expensive, so instead the single object is allocated directly in the shared young space outside any buffer, using a compare-and-swap on the shared free pointer, and the thread keeps its buffer for subsequent small allocations. HotSpot also allows only a limited number of such outside allocations before it reconsiders and refills anyway. Very large objects are never given a buffer of their own; depending on collector and size they go straight into the shared young space, the old generation, or humongous regions.
code
text · 4 lines-XX:TLABWasteTargetPercent=1 # share of allocation allowed to be wasted as tails
-XX:TLABWasteIncrement=4 # raises the limit after each outside-of-TLAB allocation
-XX:MinTLABSize=2k # floor on buffer size
-Xlog:gc+tlab=trace # per-thread refills, waste %, bytes allocated outsidego deeper
Know that the buffer eventually fills and the thread must get a new one, and that this is far rarer than allocation itself.
Describe both slow-path options and why abandoning a large tail would be wasteful.
Read buffer logging, recognize high refill counts or high outside-of-buffer fractions, and explain the contention consequence of each.
Reason about the space-versus-contention tradeoff being tuned here and why the default target is deliberately tiny and not worth touching in most systems.
## The decision at the boundary The inlined fast path ends with a comparison against the buffer's end address. When it fails, control moves into the runtime, which now has a genuine choice, and the choice matters because both options waste something. **Option A — retire and refill.** The current buffer is closed out and a new one is taken from the shared young space. Taking a new buffer is a synchronized operation (an atomic bump of the shared free pointer), which is fine because it happens once per buffer rather than once per object. The problem is the unused tail of the old buffer: if a thread has 40 KB left and asks for a fresh 256 KB buffer, those 40 KB are simply abandoned until the next young collection. That abandoned space is *TLAB waste*, and it is a form of internal fragmentation of the young space. **Option B — allocate outside the buffer.** The one oversized object is placed directly in the shared young space with a compare-and-swap on the shared pointer, and the thread's own buffer stays intact for later small objects. This wastes no space but costs a contended atomic operation. ## The refill-waste limit HotSpot picks between them with a per-thread threshold. `-XX:TLABWasteTargetPercent` (default 1) expresses how much of allocation the runtime is willing to throw away as unused tails. From it, a refill waste limit is derived: if the remaining space is *below* the limit, retiring is cheap, so refill; if it is *above* the limit, allocate outside instead. To stop a pathological loop where a thread with lots of remaining space keeps allocating outside forever, the limit is increased (by `TLABWasteIncrement`, default 4) each time an outside allocation is done, so the thread eventually refills anyway. ## Keeping the heap parsable Collectors and other runtime code need to walk the heap linearly, reading an object header, computing its size, and stepping to the next object. A retired buffer's unused tail would break that walk, so HotSpot writes a filler object over it — in practice an `int[]` of the right length, or a single small object if the gap is tiny. This is invisible to applications but explains why a heap can contain arrays nobody allocated. ## Large objects An object bigger than a whole buffer obviously cannot use the fast path. Such allocations always take the slow path. Depending on the collector and configuration, a large object may be placed in the shared young space, allocated directly in the old generation when it exceeds a size threshold, or in G1 occupy a run of contiguous humongous regions. The important interview point is simply that large-object allocation does not benefit from the fast path and involves synchronization, so allocating a stream of very large arrays behaves quite differently from allocating many small objects. ## Observing it HotSpot can report per-thread buffer statistics with `-Xlog:gc+tlab=trace` at each young collection: number of refills, bytes allocated inside buffers, bytes allocated outside them, waste as a percentage, and the buffer size chosen for each thread. Two patterns are worth recognizing. Many refills with tiny buffers suggests too many allocating threads for the young space. A high fraction of bytes allocated outside buffers suggests large objects dominating, which means the allocation path is contended and the usual reasoning about cheap allocation does not apply. ## Why it matters in practice Waste consumes young space, so it slightly increases collection frequency; it is bounded by design to about 1% of allocation and is almost never worth tuning. The genuinely actionable insights are the pathological ends: thousands of allocating threads shrink buffers and multiply refills, and workloads dominated by very large arrays skip the fast path entirely.
- Why not always allocate outside the buffer instead of retiring one and wasting the tail?Because allocating outside requires an atomic update of the shared young-space pointer, which reintroduces exactly the contention the buffers exist to remove. Doing it for every allocation would serialize allocation across threads. Wasting a small tail is the cheaper trade, which is why the choice is governed by how much space would be abandoned.
- How would you tell from a running JVM that large-object allocation is bypassing the fast path?Per-thread buffer logging shows bytes allocated outside the buffers alongside bytes allocated inside them; a high outside fraction is the signal. Allocation profiling also distinguishes samples taken on buffer refills from samples taken on outside-of-buffer allocations, so a workload dominated by big arrays shows up as the latter.
saying these in an interview costs you the question
- Claiming the leftover tail is returned to the shared free pool — it is abandoned until the next young collection.
- Saying every buffer-full event causes a garbage collection; it normally just causes a refill.
- Thinking waste is unbounded; it is targeted at roughly one percent of allocation by design.
- Assuming huge objects get their own private buffer.