skip to content

Deferred Links & Proxies

The statement that fires long after the query returned because your code touched a field, iterated a collection or called a method on a stand-in. Probed because the cost is invisible at the call site.

on this pageshow

questions

5

When a data-access layer returns a deferred stand-in for a related object, which accesses fire a load and which do not?

level: juniorimportance: must knowfreq 70%

answer

  1. invisible at the call site
  2. touch is what pays
  3. the key is already in hand
  4. size and iterate both count

basics

~20 s

Any access that needs data the row has not supplied fires the load: reading a mapped field, iterating or sizing a collection, or calling a behaviour method on the stand-in. Reading the identifier it already holds does not.

solid answer

~40 s

A deferred link is a placeholder that carries the foreign key and a way back to the layer, and nothing else. The load fires on a *touch*: reading any mapped field through an accessor, calling a method that reads one, iterating a deferred collection, or asking that collection for its size or whether it is empty. It does not fire for what the placeholder can answer alone — reading the identifier the key already carries, or checking that the reference itself is not null. Two consequences bite. Generic code that walks readable properties — logging, serialising, mapping to a transfer model — touches everything it can reach. And a touch written inside a loop fires once per iteration, so the statement count scales with the result, not with the code.

go deeper

for a junior

Remember the rule of thumb: the load happens when your code touches the data, not when the parent object came back. Reading a field, looping a collection or calling a method on the placeholder are all touches.

for a middle

Explain the trigger surface precisely: which accesses need the unread row and which the placeholder answers from the key it already carries, and why a size or emptiness question is usually not free.

for a senior

Show how you find these in a running system. Count statements per request in a test, read the emitted statement log rather than the mapping, and treat any touch inside a loop as a defect to name in the query instead.

for a principal

Frame it as a cost that is invisible at the call site. Decide where in the codebase deferred links may be touched at all, and make statements-per-request an enforced number rather than a habit.

## What "deferred" actually means When a data-access layer turns a row into an object, it does not have to bring back everything that row points at. It can read the parent row, keep the **foreign key** values it found there, and hand the caller a **stand-in** where the related object belongs — an object that looks like the target, answers to the same declared type, and carries nothing but the key and a way back to the layer. Collections get the same treatment: the field holds a real collection object, but no element rows have been read. The extra statement runs at the moment something asks for data the stand-in does not have. That moment is the **touch**. Nothing about a touch is visible at the call site — the line that pays for a network round trip looks exactly like a field read, which is the whole reason this mechanism is worth interviewing about. ## Which accesses are touches | Access | Fires a statement? | Why | |---|---|---| | Reading the identifier of a stand-in | Usually no | The key came back on the parent row | | Checking that the reference is not null | No | The stand-in itself always exists | | Reading any other mapped field | Yes | That value lives in the unread row | | Calling a behaviour method on the stand-in | Yes | It reads mapped fields to compute | | Iterating a deferred collection | Yes | The element rows must be read | | Asking a deferred collection for its size | Usually yes | Answered by a load or a count statement | | Adding an element to a deferred collection | Sometimes no | Some layers queue the addition unread | | Rendering, logging or serialising the object | Yes, repeatedly | Generic code reads every readable property | Two entries in that table are worth dwelling on. - **The identifier is free, the rest is not.** A stand-in is built around the key, so code that only needs the key — writing a reference into a message, comparing two links, building a URL — can run against an unloaded object at no cost. Code that reaches one field further pays in full. - **Emptiness is not free.** The collection reference is always non-null, so a null test proves nothing; whether there are rows is a property of the database, and answering it costs a statement. ## Why the cost is invisible Three properties conspire: 1. **The syntax is identical.** A read that hits memory and a read that opens a round trip are the same expression. Nothing in the source distinguishes them. 2. **The trigger is a runtime path, not a declaration.** Whether a link is loaded depends on what the query fetched and what already ran in this unit of work, so the same line can be free on one code path and expensive on another. 3. **Touches compound in loops.** A touch written inside an iteration over parents fires once per iteration, so the statement count scales with the result size rather than with the code. Generic code makes this worse, because it touches indiscriminately. A formatter, an object-to-object mapper or a debug rendering walks every readable property it can reach; it has no idea which links the caller intended to fetch, so it faults in whatever the graph exposes. ## Seeing it rather than guessing Deferred loads are cheap to measure and impossible to intuit: - Count the statements a code path emits and treat that count as an assertion in tests, so a new touch shows up as a failing number rather than as a slow screen months later. - Read the emitted statements, not the configuration: a mapping that says "deferred" proves intent, and only the statement log proves behaviour. - Look at *when* each statement fires relative to the parent query. A cluster of small, identical statements after one larger one is the signature of touches. ## Working with it deliberately The mechanism is not a defect — deferring is what keeps a read of one object from dragging back an unbounded graph. The discipline is to make the touch a decision instead of an accident: - Decide per use case which links the query should bring back, and let the rest stay deferred. - Keep touches out of loops; if a loop needs a link, the link belonged in the query that fed it. - Keep generic traversal code away from mapped objects, or guard it, so that formatting output never becomes a reason to query. - Remember that layers differ: some defer references by default and collections always, some load everything unless told otherwise, and a query builder with no tracked objects has no deferral at all — it returns rows, and a missing join is simply missing data. Knowing which kind of layer you are in tells you whether an absent value means "not fetched yet" or "not fetched, ever".

  • Does adding an element to a deferred collection always force it to load first?
    Not necessarily. Several layers record the addition against the deferred collection and leave it unread, because appending a row does not require knowing the existing ones. Others load first, and any collection that must reject duplicates has to load in order to check. Treat it as layer-specific and confirm with a statement count rather than an assumption.
  • Why does asking a deferred collection whether it is empty usually cost a statement, when the reference is clearly non-null?
    The reference is the layer's own collection object, which always exists; emptiness is a property of the rows, which have not been read. Layers answer it either by loading the elements or by running a count. Both are statements, so an emptiness guard written to avoid work is itself work.

A stand-in is a claim ticket. Holding it and reading the number costs nothing; asking what is in the parcel sends someone to the storeroom.

saying these in an interview costs you the question

  • Thinks the related rows already came back with the parent and are merely revealed later
  • Believes asking a deferred collection for its size or emptiness is free
  • Assumes reading the identifier off a stand-in must fetch the whole row
  • Treats logging or serialising an object as a read-only observation
  • Expects the load to fire when the method ends rather than at the touch
open as a page

Why does a lazily loaded stand-in break a type test, a cast, or an equality comparison against the mapped class?

level: middleimportance: must knowfreq 58%

basics

~20 s

A stand-in is an instance of a type the layer generated, not of the mapped class. Exact-class tests fail, casts to a concrete subtype fail even after loading, and equality over runtime classes is asymmetric.

open as a page

Why does a mapper replace the collection instance a constructor assigned with its own instrumented collection?

level: middleimportance: should knowfreq 44%

basics

~20 s

The layer needs a collection it controls: one recording whether the rows were read, which object owns it, and what was added or removed. On load it installs its own implementation, discarding the constructor's instance.

open as a page

How do you check whether a deferred link is already loaded before touching it, and when is that check the right tool?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Ask the layer, not the object: a load-state predicate reports whether a link or one attribute has been fetched, without touching it. Use it in generic code that walks object graphs and cannot know the caller's fetch plan.

open as a page