skip to content

Fetching Without Blocking

What changes when no thread may wait on a statement: nothing faults in on access, so every link is named up front or fetched by a second call. Asked to see why the stand-in disappears entirely.

on this pageshow

questions

4

In a non-blocking data-access layer, why can a deferred association not fault in on access, and what replaces it?

level: middleimportance: must knowfreq 52%

answer

  1. nowhere to wait
  2. an accessor must return immediately
  3. named in the query, or a second call
  4. forgetting shows as empty, not as an error

basics

~20 s

A stand-in faults in by making the caller wait, and a non-blocking layer has nowhere to wait. So a plain accessor cannot load: links are named in the query, exposed as a deferred handle, or fetched by a second call.

solid answer

~50 s

Transparent lazy loading rests on one prerequisite: the thread touching the field is allowed to stop until rows come back. A layer built so that no operation parks a thread cannot honour that, because a plain accessor has to return something immediately. So the stand-in mechanism largely disappears. What replaces it is explicit: name the link in the query that loads the parent (a join or fetch plan), collect the parent keys and issue **one composed second statement** for the children, or read a projection with exactly the columns the caller needs. Some layers keep a link visible in the type system as a pending value the caller must resolve at a suspension point, which is still explicit. The important consequence is the failure mode: forgetting a link no longer throws when you touch it, it usually shows up as an absent reference or an empty collection in the response.

go deeper

for a junior

Remember the one-line reason: loading on access needs the caller to wait, and here nothing may wait, so what you did not ask for is simply not there.

for a middle

Explain the mechanism both ways: why a placeholder cannot issue a statement inside a plain accessor, and the three explicit replacements — name it in the query, fetch it with one composed second call, or project the columns you need.

for a senior

Talk about the failure mode you now own. Forgotten links return empty data instead of raising, so you argue for content assertions, statement counting, and a representation that distinguishes not-fetched from not-present.

for a principal

Frame it as moving a decision, not deleting one. Fetch shape becomes part of every read path's contract; weigh the reviewability that buys against the cost of losing a mapper convention that used to cover forgotten cases.

## What faulting in actually requires A **deferred link** is a placeholder the mapper puts where an association should be: a stand-in object, or an instrumented collection. The first time application code touches it, the placeholder issues a statement, waits for the rows, populates itself, and returns the value. The whole appeal is transparency — reading `order.customer` looks like ordinary memory access even though it is a round trip to the database. That trick has one hard prerequisite. The thread executing the accessor must be allowed to **stop and wait** for the statement to finish. A layer designed so that no operation ever parks a thread cannot grant that: the thread must return control and be resumed later, when the rows arrive, through a continuation. An ordinary accessor has no way to suspend halfway through returning a value, so it must hand back whatever is already in memory. Hence the rule of thumb: **where nothing may block, nothing faults in.** ## What takes the stand-in's place - **Name the link in the query that loads the parent.** A join, or a declared fetch plan attached to the read, so the association arrives with the root in one round trip. - **Issue a composed second call.** Load the parents, take their keys, run *one* statement for all matching children, and stitch the graph in memory. Composed means one statement for the whole batch, not one per parent. - **Project instead of navigate.** Read a transfer shape containing exactly the columns the caller needs, so there is no graph to walk in the first place. - **Keep the link explicit in the type.** Some layers model an unfetched association as a pending value that the caller has to resolve at a suspension point. The load still happens, but it is visible in the signature rather than hidden behind a getter. | Concern | Layer where blocking is allowed | Layer where it is not | |---|---|---| | Touching an unfetched link | a statement fires transparently | no statement can fire | | Where fetching is decided | at the point of use, implicitly | in the query or the call chain, explicitly | | Cost of forgetting a link | a hidden extra statement, or a failure once the unit of work has closed | usually missing or empty data in the answer | | How you notice | an error at the touch site, or statement counters | tests asserting content, plus statement counters | ## The failure mode moves from loud to quiet This is the part that surprises people porting a read path. In a blocking mapper, an association you forgot to fetch still produces the right answer — slowly — and, if the unit of work has already closed, it usually produces an error that points straight at the line. Neither safety net exists here. An unfetched reference is commonly absent, and an unfetched collection is commonly empty, so the response is **wrong rather than broken**: a customer that appears to have no orders, a total computed over nothing. Practical defences: 1. Assert on content, not on absence of exceptions — a test that only checks the call succeeded will pass on an empty graph. 2. Prefer a representation that keeps the foreign key visible (a key-only stub) over one that yields nothing, so "not fetched" is distinguishable from "no such row". 3. Never treat an empty collection as evidence that the database holds no children, unless the query that produced it demonstrably fetched them. ## The repeated-statement trap does not vanish Making fetching explicit does not by itself make it efficient. A loop that resolves each parent's link with its own call reproduces exactly the one-statement-per-row pattern, now hand-written. Fanning those calls out concurrently makes it worse in a specific way: instead of N statements in sequence, you get a burst of N statements competing for a bounded pool of connections, which can stall the whole service rather than just this request. The composed second call — one statement keyed by all the parent identifiers — is the shape to reach for. ## What does not change The database side is untouched. Row-to-object mapping, de-duplicating a joined result, the tradeoff between one wide join and two narrower statements, transaction boundaries, and the cost of a round trip are all the same problems with the same answers. Removing the stand-in changes **who decides what is loaded, and when** — it moves that decision from the point of use to the point of query — and it removes the implicit safety net that used to cover a forgotten decision.

  • If unfetched links are silently empty, how do you keep a read path honest as it changes?
    Assert on content in tests for every association the response is supposed to carry, and count statements per code path so a hand-written per-parent loop shows up as a regression. Where the layer can represent an unfetched link as a key-only stub rather than nothing, use it: absence then means "no row", not "never asked".
  • Does the disappearance of stand-ins remove the case for loading part of a graph later?
    No. Deferring is still right when the extra data is needed on only a few requests. What changes is that deferral becomes a second explicit call the caller writes and can batch, rather than something that happens behind an accessor. The decision survives; the transparency does not.
  • Is an association fetched by a second statement equivalent to one fetched by a join?
    Not exactly. Two statements can observe different states unless both run inside one transaction with a suitable isolation level, and they cost two round trips instead of one. In exchange they avoid a joined result multiplying parent rows. The choice is the usual fetch-plan tradeoff, just made explicitly.

A vending machine can restock a slot while you stand there; a machine that is forbidden to make you wait can only hand you what is already in the slot, so someone has to load it in advance.

saying these in an interview costs you the question

  • Claims lazy loading still works, just asynchronously behind the same getter
  • Assumes a forgotten association throws, so the mistake will be obvious
  • Reads an empty collection as proof the parent has no children
  • Resolves each parent's link with its own call and calls that explicit fetching
  • Says briefly blocking on the pending value is fine because the query is fast
  • Thinks removing stand-ins also removes the need to decide what to fetch
open as a page

When a non-blocking read hands back a stream of rows instead of a materialised list, what changes for the caller and the connection?

level: middleimportance: should knowfreq 44%

basics

~20 s

The call returns before any statement runs: work starts when the caller consumes, rows arrive one by one, and errors can surface mid-sequence. The connection stays checked out for the whole consumption, so a slow consumer holds it.

open as a page

In a non-blocking service, how can a lost asynchronous context make statements run outside the intended transaction or on a second connection?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A blocking layer keeps the current unit of work in thread-bound storage. Non-blocking work resumes on other threads, so that slot is empty or holds another request's context, and the statement quietly runs on its own connection.

open as a page

Your database work already runs behind a bounded connection pool, so what does a non-blocking data-access layer buy, and when is keeping the blocking one better?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

The pool bounds database concurrency either way, so a non-blocking layer buys no extra throughput. It buys threads: waiting requests stop occupying them, so overload queues instead of exhausting a thread pool. It costs explicit fetching and manual context.

open as a page