How do pre-allocated key blocks let a data-access layer know an object's identifier before the row is written?
answer
- one round trip buys many keys
- the counter moves in its own transaction
- values handed out from memory
- the remainder of a block is discarded
- block size trades trips against waste
basics
~20 sThe layer claims a whole range from a store-side counter in one round trip, then hands values out from memory. The identifier therefore exists the moment the object is constructed; whatever is left in the range when the process stops is discarded.
solid answer
~40 sRather than letting each insert produce its own value, the layer reserves a block: it advances a shared counter by a fixed amount in one short, self-contained transaction and keeps the resulting range in memory. Objects then take identifiers from that range with no round trip at all, so the key is set at construction — before the write, before the flush, before a transaction even has to exist. Two processes hold disjoint ranges, so their numbers interleave rather than collide, and values are no longer in global insert order. The cost is discarded values: a restart, a rollback, or a short burst leaves the remainder of the range unused. Block size is the knob — larger blocks mean fewer counter round trips and larger discarded ranges.
go deeper
Learn the shape of it: one trip to a shared counter reserves many identifiers at once, so an object gets its key before anything is written to the store.
Be able to explain the mechanics — the counter advances by the block size in its own short transaction, the range is served from memory, and the remainder is discarded whenever the process stops.
Bring the operational consequences: interleaved numbering across instances, a range lost on every deploy, and the sizing trade-off between counter contention and how fast the number space is consumed.
Argue the strategy rather than the number. Buying keys ahead takes the store out of the write path's critical section, at the price of identifiers that are neither dense nor globally ordered; decide first whether anything downstream depends on those properties.
**Pre-allocated key blocks** move the moment a key becomes known from the insert to the constructor. Instead of asking the store for one value per row, the layer asks once for many, and then serves identifiers out of memory until the reservation is used up. ## The mechanism A shared counter lives in the store — a counter object, or a row in a small allocation table that is read and incremented under a lock. When the layer needs identifiers it does one operation: advance that counter by the **block size**, and take the range between the old and new values as its own. From then on: - every new object takes the next value from the in-memory range, with **no round trip**; - when the range is exhausted, the layer advances the counter again and takes the next range; - no other process will ever be handed a value from a range already taken, because the counter never goes backwards. The reservation is a claim on a slice of the number space, not a promise that the slice will be used. ## Why the allocation stands outside your transaction The counter is advanced in its **own short transaction**, committed immediately, independent of whatever work asked for it. This is not an optimisation, it is a correctness requirement. If the increment participated in the caller's transaction, a rollback would undo it, and the same range could then be handed to a second process — two objects, one identifier, discovered later as a duplicate-key failure or, worse, as a silent overwrite. Making the allocation independent means a rolled-back or abandoned block is *wasted*, never *recycled*, and that is the trade the strategy accepts. ## What the application gains 1. The identifier exists at construction, so the object can be referenced immediately — put into the working set, used as a map key, handed to another component. 2. Both ends of a link can be wired before anything is written, because the child can carry the parent's key value already. 3. An identifier can be returned to a caller, logged, or attached to an outgoing message before commit. 4. The layer never has to read a value back per row, so the write path is a pure series of statements the layer already knows the content of. ## What it costs - **Sparse values.** Every stopped process discards the tail of its block. Deployments, crashes and scale-downs each leave a hole the size of the remainder. - **No global ordering.** Two instances holding different ranges write interleaved numbers, so a higher identifier does not mean a later row. - **Faster consumption of the number space.** With large blocks the counter climbs far ahead of the row count; a narrow numeric column reaches its ceiling sooner than the row count suggests. - **A shared point remains.** The counter is still contended, just far less often. ## Sizing the block | block size | counter round trips | values discarded per stop | numbering | |---|---|---|---| | small (tens) | frequent, one per few rows | few | close to dense | | large (thousands) | rare, amortised over many rows | up to a whole block | visibly sparse | The right size follows the write rate and the restart rate, not a default: a process that inserts thousands of rows per second and restarts weekly wants a large block; a process that writes a handful of rows a day and restarts on every deploy will burn most of a large block for nothing. ## Where layers differ Data-access layers differ in what they reserve and how widely they share it: some keep one counter for the whole model, some one per mapped type, and some combine a reserved high part with a locally incremented low part so that a single stored number stands for a whole range. The externally visible behaviour is the same in every case — the key is known before the write, and the values are not dense. ## The judgement call Pre-allocation is the right answer when the write path needs identifiers early or writes in volume, and the wrong answer when something downstream has quietly come to depend on identifiers being dense, ordered, or a proxy for a row count. Those dependencies are usually undocumented, which is why changing an existing model's key timing is a bigger change than it looks.
- Why must the counter be advanced in a transaction separate from the caller's?Because the reservation has to survive the caller's rollback. If the increment rolled back with the work, the same range could be handed to another process and two objects would claim one identifier. Independence means a discarded block is wasted, never reused.
- How does the block size change behaviour under load?A larger block cuts counter round trips roughly by that factor and reduces contention on the shared counter, but every stopped process discards up to a whole block and the number space is consumed faster. Small blocks give denser values at the price of frequent trips.
- What does the application actually gain from knowing the key before the write?The object can be referenced and linked immediately, an identifier can be returned or published before commit, and the layer never reads a value back per row. The write becomes a statement whose content the layer already knows in full.
saying these in an interview costs you the question
- Advances the shared counter inside the caller's transaction
- Treats discarded values as a bug to fix by resetting the counter
- Expects identifiers to stay ordered across several instances
- Assumes a larger block costs nothing because it is in memory
- Says two processes can share one block by taking turns