skip to content

Object Allocation & TLAB

On the fast path, allocation is just bumping a pointer inside a thread-local buffer, so no synchronization is needed until that buffer runs out. Interviewers use it to explain why object allocation on the JVM is cheap, and why allocation rate rather than object count drives collection frequency.

on this pageshow

questions

5

What is a thread-local allocation buffer (TLAB) in the HotSpot JVM, and why does having one make object allocation cheap?

level: middleimportance: must knowfreq 60%

answer

  1. Eden empty + contiguous → bump the pointer
  2. shared top = CAS + cache-line ping-pong
  3. three words: start / top / end, thread-private
  4. fast path ~10 instructions, no lock, no fence
  5. filler object keeps heap parsable

basics

~20 s

A TLAB is a private chunk of the young space handed to one thread. Inside it the thread allocates by advancing a pointer, with no lock and no atomic instruction. Only buffer refills and oversized objects touch the shared heap.

solid answer

~50 s

HotSpot allocates most objects in Eden, the young space, which after a young collection is one empty contiguous range. Allocating there is just advance a free pointer by the object size — bump-the-pointer. The catch is that the pointer is shared: if every thread bumped the same word, each allocation would need an atomic compare-and-swap, and that cache line would bounce between cores. A TLAB removes the sharing. Each Java thread gets its own slice of Eden, tracked by three fields in the thread structure: start, current top, and end. JIT-compiled allocation becomes roughly: load top, add size, compare against end, store the new top, write the object header, return the reference — about a dozen instructions, no synchronization, because no other thread can allocate from that buffer. When the buffer runs out the thread takes a slow path: retire it and request a fresh one from the shared heap (synchronized, but once per many allocations), or allocate the object directly outside any buffer.

code

text · 4 lines
text
-XX:+UseTLAB            # on by default; -XX:-UseTLAB is for debugging only
-XX:TLABSize=256k       # fixed starting size instead of the adaptive one
-XX:-ResizeTLAB         # stop adapting size per thread
-Xlog:gc+tlab=trace     # per-thread TLAB fill/refill/waste statistics

go deeper

for a junior

Be able to say allocation is a pointer bump in a per-thread chunk of the young space, which is why creating small short-lived objects is cheap.

for a middle

Explain the three fields (start/top/end), why the shared alternative needs an atomic operation, and what triggers the slow path.

for a senior

Connect it to observed behaviour: allocation throughput scaling across cores, zeroing cost, locality effects, and how thread count divides Eden.

for a principal

Frame the design tradeoff — partitioning removes contention at the cost of internal fragmentation and per-thread footprint, and reason about when that tradeoff bends (huge thread counts, tiny heaps).

## The shape of the young space Almost every `new` on HotSpot lands in the young part of the heap, conventionally called Eden. Every young collector in current HotSpot moves objects: it copies the small live set elsewhere and leaves the collected space empty. Because of that, free memory in Eden is not a linked list of holes — it is one contiguous range described by a start and an end address. Allocating N bytes is then arithmetic: take the current free pointer, add N, and if the result is still below the end, the object occupies the old pointer's address. This is called bump-the-pointer allocation, and it is a handful of instructions rather than the search-a-free-list work a general allocator such as C's malloc must do. ## Why one shared pointer does not scale If all application threads bumped one shared pointer, the read-modify-write would have to be atomic, typically a compare-and-swap loop. Two costs follow. First, under contention the CAS fails and retries. Second, and worse, the cache line holding the pointer is written by every allocating thread, so it migrates between cores continuously; on a 32-core machine allocating hard, that single line becomes the bottleneck and allocation throughput stops scaling with cores. ## The buffer A TLAB fixes this by partitioning. Each Java thread carves out a slice of Eden for its exclusive use — typically tens to hundreds of kilobytes — and keeps three words in its own thread structure: the buffer start, the current top, and the end. Because only the owner thread ever reads or writes its own top, no atomicity is required at all. The JIT inlines the fast path directly into compiled code: 1. load `top` from the thread structure; 2. add the object's size (known statically for a class, computed for an array); 3. compare with `end`; if it exceeds it, jump to the slow path; 4. store the new `top`; 5. initialize the object header (mark word plus class pointer) and any fields not covered by pre-zeroed memory; 6. return the old `top` as the reference. There is no lock, no CAS, and no memory fence on this path. The claim that JVM allocation is faster than C malloc is about exactly this sequence. ## What still costs something Allocation is cheap, not free. Java requires fields to start at their default values, so the memory must be zeroed; HotSpot either zeroes the whole TLAB when it is handed out or zeroes each object as it is allocated, and either way the cost is proportional to bytes allocated. Header initialization is a couple of stores. And every allocated byte is future work for the collector: bytes allocated per second determine how often the young space fills and therefore how often young collections run. A pleasant side effect of thread-local buffers is locality. Objects allocated close together in time by one thread are contiguous in that thread's buffer, so they tend to share cache lines and are prefetch-friendly, and two threads' fresh objects never share a cache line, which avoids accidental false sharing between unrelated threads. ## When the fast path is not taken Three cases leave the fast path: the object does not fit in the remaining buffer (refill or allocate outside the buffer), the object is large enough that giving it a buffer would be wasteful (allocated directly in the shared young space, or straight into the old generation or a humongous region depending on collector and size), and TLABs being switched off with `-XX:-UseTLAB`, which is a debugging option only — doing it in production collapses allocation onto a synchronized shared path. One invariant matters for correctness: the heap must remain parsable, meaning a collector can walk it object by object. Since a retired buffer usually has unused tail space, the JVM fills that gap with a dummy filler object (typically an int array) so the walk does not run into uninitialized memory.

  • If the buffer is thread-private, do objects allocated in it stay private to that thread?
    No. The buffer is private only for the purpose of allocation. The reference can be published to any other thread immediately, and other threads read and write the object normally. Thread-locality here is an allocation-time property of the memory range, not a visibility or ownership property of the object.
  • Does having many threads make TLABs less effective?
    It can. Eden is divided among the threads that allocate, so with thousands of threads each buffer becomes small, refills become more frequent, and unused tails of retired buffers add up as wasted Eden. Threads that allocate very little still get a minimum-size buffer, so a large pool of mostly idle threads can consume young space without doing useful allocation.

A shared pointer is one cashier for the whole store; a TLAB is giving each shopper a pre-counted roll of coins so they only queue when the roll runs out.

saying these in an interview costs you the question

  • Claiming a TLAB is a separate memory region outside the heap — it is a slice of Eden inside the normal heap.
  • Saying allocation uses a compare-and-swap or a lock on the fast path; the whole point is that it needs neither.
  • Assuming objects allocated in a thread's buffer are somehow private or thread-confined.
  • Believing allocation is literally free — zeroing costs bytes and every allocation is future collector work.

context

open as a page

Developers often say allocating a small object on the JVM is cheaper than calling malloc in C. Mechanically, what does the JVM do on allocation that makes that plausible, and what does it pay for it later?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Because a moving young collector leaves free space contiguous, allocation is just advancing a pointer and writing the object header, with no free-list search. The bill arrives later: the collector must find and copy the surviving objects.

open as a page

A service allocates roughly 1 GB per second of mostly short-lived objects and its young space is 512 MB. How does allocation rate translate into collection frequency, and what levers change that?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Young collections happen roughly every time the young space fills: 1 GB/s into 512 MB is about two young collections per second. Levers are allocating fewer bytes, enlarging the young space, and reducing how much survives each cycle.

open as a page

A thread's private allocation buffer in the Java heap no longer has room for the next object. What choices does the HotSpot JVM make at that point, and what is meant by allocation-buffer waste?

level: seniorimportance: should knowfreq 30%

basics

~20 s

It either retires the buffer and grabs a fresh one, or allocates that one object in the shared young space. Retiring abandons the unused tail, which is waste; HotSpot only retires when the remaining tail is below a threshold, otherwise the big object goes outside the buffer.

open as a page

How does HotSpot decide how large each thread's private allocation buffer should be, and when, if ever, would you override that decision with tuning flags?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

HotSpot sizes each buffer adaptively from that thread's recent allocation history and the young space available per thread, resizing at each collection. Overriding is rare; the real cases are pathological thread counts, tiny heaps, or workloads where logging shows excessive refills.

open as a page