skip to content

A data-access interface hides how its results are fetched. Which fetching decisions can it hide from callers, and which cannot?

level: seniorimportance: must knowfreq 58%

answer

  1. mechanism hides, cost does not
  2. statements, volume, timing, tracking
  3. the caller pays the latency
  4. put the cost in the contract

basics

~20 s

It can hide the statement text, join strategy, chosen columns and index or cache use. It cannot hide cost or timing: statement count, what was left unloaded, whether the result is tracked, volume, paging stability and scope duration.

solid answer

~50 s

The mechanism is hideable; the cost is not. Behind a named method the layer can change the statement text, the join strategy, which columns it reads, whether a cache answered, and how it maps rows — callers never see it, and that freedom is the point of the boundary. What reaches callers regardless is how many statements the call runs, what came back loaded versus left out, whether the result is tracked, how many rows it holds, whether a paged ordering is stable, and how long a scope has to stay open. Each of those shows up as latency, a later query, a failed read, a lost update or a lock held too long. So the contract has to state them: name what is loaded, make tracked-versus-snapshot visible in the return type, pin paging orders, and assert statement counts in tests rather than describing them in comments.

go deeper

for a junior

Remember the split: an interface can change how a query is written without touching callers, but it cannot make the query free. Cost and timing still reach whoever called it.

for a middle

List what leaks and how each one shows up - statement count as latency, unloaded parts as a later query or a failure, tracking as whether a mutation gets written, volume as memory.

for a senior

Demonstrate that you turn those leaks into contract: distinct methods for distinct fetch shapes, return types that reveal tracking, pinned paging orders, and statement-count assertions in tests.

for a principal

Take a position on how much the boundary should promise. Over-promising invites callers to ignore cost; over-specifying freezes the layer. Say where you draw it and how the team keeps it there.

## What a boundary genuinely hides A data-access interface exists partly to keep query mechanics out of its callers, and it does that well. Behind a named method the layer is free to change: - the **statement text** and the join strategy used to assemble the result; - **which columns** are read, and whether a value is computed in the database or afterwards; - whether an index, a materialised structure or a shared cache answered part of the request; - how **keys are generated** and how results are mapped onto objects; - whether a filter is one predicate or several, as long as the result set is the same. None of that is observable from the call site, and all of it can be improved without touching a single caller. That is the real payoff of the boundary, and it is worth defending. ## What escapes anyway What the boundary cannot hide is the **cost and the timing** of the fetch, because the caller experiences both. 1. **How many statements the call runs.** A method that looks like one lookup may run one statement per row of a collection it walked. The caller sees it as latency, the database sees it as load, and no amount of naming hides it. 2. **What was loaded and what was left out.** If part of the graph is unloaded, the caller finds out by touching it: either another query runs, or - outside the scope that owned the object - the read cannot be satisfied at all. Whether the caller must avoid touching that part is a fact about the returned value, so it belongs in the contract. 3. **Whether the result is tracked.** A caller that mutates a tracked object writes to the database; a caller that mutates a snapshot writes nothing. Same-looking code, opposite outcome. 4. **Volume.** A method that returns everything matching a filter has handed the caller a memory profile that depends on data, not on code. 5. **Ordering and paging stability.** A page defined by an offset over an unstable ordering skips and repeats rows under concurrent writes. Callers build on the ordering whether or not the contract mentions it. 6. **Transaction and lock scope.** Streaming a large result keeps a scope, and often a snapshot or locks, alive for as long as the caller takes to consume it. The caller controls that duration. | Decision behind the boundary | Hideable? | How it reaches the caller | |---|---|---| | Statement text and join strategy | Yes | Not observable | | Which index or cache answered | Yes | Not observable | | Number of statements per call | No | Latency and database load | | Loaded versus deferred parts | No | Later query, or a failed read | | Tracked versus snapshot result | No | Whether a mutation is written | | Result volume | No | Memory, and time to first use | | Ordering and paging stability | No | Skipped or repeated rows | | How long a scope stays open | No | Lock and connection hold time | ## The rule this leads to A boundary can hide the **mechanism** and not the **cost**. So the contract has to state the cost - not in prose nobody reads, but in the shape of the interface: 1. **Name what comes back loaded.** A method whose result is safe to walk in full is a different method from one whose result is a stub, and they should not share a name. 2. **Make the tracked-versus-snapshot distinction visible in the return type**, so a caller can see whether mutating the result means anything. 3. **Separate the streamed or unbounded read.** A method that requires an open scope and consumes as it goes should say so, because the caller now owns the duration. 4. **Pin the ordering** of anything paginated, on a key that does not move. 5. **Assert the statement count** of important call paths in an automated test. Prose about how many queries a method runs decays; an assertion fails on the commit that changed it. ## Why "add another method" is not always the fix When a caller needs a differently-fetched result, the tempting move is a second method with a longer name, and then a third. That is often right - it keeps the fetch decision inside the layer, where it can be seen and changed. It stops being right when the methods multiply because callers are really composing arbitrary filters through the boundary; at that point the interface is a query language with an awkward syntax, and the honest options are a named read boundary per use case or an explicit query type. Layers also differ in how much of this is even negotiable: some resolve unloaded parts on access, which turns a fetch decision into a latency surprise, while others require every part to be requested up front and fail loudly instead. The first hides more and leaks later; the second hides less and leaks immediately. Knowing which behaviour a layer has tells you how much the contract has to say out loud - and neither behaviour lets the boundary claim that fetching is somebody else's problem.

  • Why is a comment describing how many statements a method runs a weak contract?
    Because nothing enforces it. A change to the fetch plan, a new association, or a caller that walks one more collection can multiply the count, and the comment still reads the same. An automated assertion on the statement count of the call path fails on the commit that broke it, which is the only feedback that survives turnover.
  • How does an unstable paging order leak through a boundary that never mentions ordering?
    Callers page by offset and assume rows keep their positions. Under concurrent inserts and updates an unpinned ordering shifts rows between pages, so some are shown twice and others never appear. The boundary said nothing, but the caller built on the behaviour, so it is part of the contract in practice.
  • If a method returns a partially loaded graph, what should the interface do about it?
    Say so in the method, not in documentation. A result that is safe to walk in full and a result that is a stub are different contracts and deserve different methods. Anything else leaves each caller to discover the boundary of what was loaded by touching it and seeing what happens.

saying these in an interview costs you the question

  • Claims a good interface makes fetching entirely the layer's problem
  • Thinks statement count is invisible to callers because they see only the result
  • Assumes callers cannot tell whether a returned object is tracked
  • Believes documentation is enough to pin a method's query count
  • Pages by offset over an ordering the contract never fixes
  • Treats a streamed read as ordinary, ignoring who holds the scope open