skip to content

Key Generation & Timing

When a key becomes known: assigned at insert, drawn ahead in pre-allocated blocks, or computed by the application. Asked because code that needs the key early pays for the wrong choice at every write.

on this pageshow

questions

5

When the store assigns a key during the insert, at what point does a mapped object learn its identifier?

level: juniorimportance: must knowfreq 72%

answer

  1. the store decides, not the mapper
  2. empty until the row is written
  3. flush emits the insert first
  4. value read back from the statement
  5. asking early forces one write

basics

~20 s

Only when the insert actually runs. Until that statement reaches the store the key field is empty, so asking a pending object for its identifier forces the layer to write that row immediately and read the generated value back.

solid answer

~40 s

A store-assigned key is produced by the database as the row is written, so it cannot exist in the object beforehand. When you hand a new object to the layer it joins the unit of work as a pending insert with an empty key, and the statement is normally deferred until flush. At flush the layer emits the insert and reads the generated value back — some engines return it with the statement, otherwise the connection exposes the value it just produced — then writes it onto the object and files the object in the identity map under it. Touching the identifier before that point forces an early, single-row flush, which is why code that needs the key up front pays for it at every write.

go deeper

for a junior

Remember the order: register, flush, insert, then the key exists. Before the insert has run the key field is empty, and any code reading it is reading nothing useful.

for a middle

Explain that the statement is deferred to flush, that the generated value is read back through the channel the statement provides, and that the object is only then filed in the identity map under its key.

for a senior

Show where this bites in production: a path that needs the identifier to return a location, name a file, or publish a message pays a round trip per row, and a bulk write stops behaving like one.

for a principal

Frame it as a cost curve rather than a rule. Buying the identifier earlier trades counter waste and loss of insert ordering for a write path that does not have to consult the store before it can describe its own data.

A **store-assigned key** is a value the database produces at the moment the row is written: the column is declared so that the engine supplies a number when the insert does not name one. The application never chooses it. That single fact drives everything else here — a value the store creates cannot exist in memory before the store has been asked to create it. ## Why the object sits around with an empty key Most mapping layers do not send an insert the instant you hand them a new object. The object joins the **unit of work** — the set of changes the layer is accumulating — as a pending insert, and the statement is deferred to the next **flush**: an explicit call, an automatic flush before a query whose results the pending row could change, or the flush that runs just before commit. Between registration and flush the object is perfectly usable in memory and its key field holds nothing: a null reference, or a sentinel such as zero for a numeric field that cannot be null. ## The sequence, step by step 1. Your code constructs the object. The key field is empty. 2. You register it with the layer. It is recorded as new, but it is *not* filed in the **identity map** — the layer's one-instance-per-row index — because that index is keyed by identifier and there is no identifier yet. A separate set of new objects holds it, keyed by object reference. 3. The layer flushes. The insert is emitted without naming the key column, so the engine supplies the value. 4. The layer reads the value back. Some engines hand it back as part of the same statement; otherwise the connection exposes the value the statement just generated. Either way the read-back is *per statement*. 5. The layer writes the value into the key field and files the object in the identity map under it. Only after step 5 does the object's `key` mean anything to you, to a log line, or to another system. ## What makes the write happen sooner than the layer wanted - **Reading the identifier at all** — the layer must produce a real value, so it flushes that row now. - **Returning a location or an identifier to a caller** before the transaction commits. - **Storing a link as a key value** rather than as an object reference: the child row cannot be written until the parent's key exists. - **Publishing a message, naming a file, or writing an external artefact** after the row. - **Putting the object into a hash-based collection** whose hash reads the key, which forces the write just to get a stable value. Each of those converts a deferred write into a statement sent right now. This is why the timing of the key quietly shapes the whole write path, and — as a consequence only — why writes that could otherwise have travelled together stop travelling together. ## Three timings, compared | when the key is known | what produces it | extra round trip | usable at construction | |---|---|---|---| | at insert | the store, while writing the row | a read-back per statement | no | | ahead of time | a store-side counter, drawn in blocks | one per block, amortised | yes | | at construction | the application computes it | none | yes | The first row is the subject of this question; the other two exist precisely because the first one makes an early identifier expensive. ## The shortcut that is not one A recurring mistake is to write the row and then read the key with a query for the largest value in the key column. It is wrong for a reason that testing rarely exposes: another session inserting at the same time advances the same counter, so the largest value may belong to someone else's row. The value must come from the statement that produced it, through the channel the layer already uses, not from a second query. ## Where layers differ Data-access layers differ here, and it is worth knowing which kind you are in. Some flush automatically before any query whose results a pending row could change, so by the time you look the insert has usually already happened; others write only when told, and the same code sees an empty key for much longer. Some always read the generated value back; others let you skip the read-back deliberately when you do not need the key. Do not carry one layer's timing into your reasoning about another. ## What to take away - The empty key is not a bug; it is the honest state of an object whose row does not exist yet. - The read-back is the layer's only way to learn the value, and it happens per insert. - Any requirement for the identifier before the write is a requirement to change *when* the key is produced, not a requirement to read harder.

  • Which situations force the layer to send a pending insert earlier than it intended?
    Reading the identifier, wiring a child row whose link is stored as a key value, returning an identifier to a caller, and any query whose results the pending row could change if the layer flushes automatically. Each of them needs the row, or its key, to be real now.
  • How does the layer keep track of a new object before it has a key?
    By object reference, in a separate set of new objects. The identity map is keyed by identifier, so an object without one cannot be filed there; it moves into the identity map after the insert returns a value.
  • What should code do when it genuinely needs the identifier before commit?
    Either accept an early flush of that one row inside the same transaction — the value is visible to you and rolls back with you — or move the key off the store entirely by drawing it from a pre-allocated block or computing it in the application, so it exists at construction.

saying these in an interview costs you the question

  • Thinks the layer picks the value and the store merely records it
  • Expects the identifier to be filled as soon as the object is registered
  • Assumes a deferred insert already ran because no error appeared
  • Fetches the new key by selecting the largest value in the column
  • Passes a pending object's empty key to another system
open as a page

What does a mapper demand of a composite key class used as a mapped object's identifier?

level: middleimportance: should knowfreq 44%

basics

~20 s

That it behave as a value: all parts populated, equality and hashing over every part, and no mutation once the row exists. The layer files objects under that whole value, so a changed key loses the row it addressed.

open as a page

How do pre-allocated key blocks let a data-access layer know an object's identifier before the row is written?

level: middleimportance: should knowfreq 55%

basics

~20 s

The layer claims a whole range from a store-side counter in one round trip, then hands values out from memory. The identifier therefore exists the moment the object is constructed; whatever is left in the range when the process stops is discarded.

open as a page

How should a mapped class define equality and hashing when its identifier stays empty until the row is written?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Never let the hash depend on a key that is filled in later: it changes when the row is written and strands the object inside any hash container it already joined. Keep the hash stable, and compare identifiers only when both exist.

open as a page

How does a data-access layer decide whether a saved object needs an insert or an update, and what breaks when the application assigns keys?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Usually by a marker meaning never-written: an empty identifier, an absent version, or the layer's own tracking. An application-assigned key is populated from birth, so the layer reads a new object as existing and updates a row that never existed.

open as a page