What does query-first, or 'design for the read', modelling mean in a document database?
answer
- Start from behaviour, not from nouns
- Write down the operations first
- Shape is a physical commitment
- Rank by frequency and latency budget
- Check the write cost of every read shape
basics
~20 sQuery-first modelling starts from the list of operations the application must serve and shapes documents so each one is answered by a simple, well-targeted access. The document shape follows the workload, not a diagram of entities.
solid answer
~50 sQuery-first modelling inverts the usual order: instead of drawing entities and then figuring out how to query them, you enumerate the application's actual access patterns — what is read, by what key, how often, and what shape the caller needs — and design documents so the hot ones are served by a single lookup. Concretely you write down each operation ("render the order detail page", "list a customer's last 20 orders"), note its frequency and latency budget, and then ask what one document would have to contain to satisfy it. The shape that answers the most and hottest operations wins, and the rarer patterns pay the cost of extra work. The point is not that reads matter more than writes — it is that in a document store the shape is a physical commitment, so it should be made against known behaviour rather than against a taxonomy.
code
json · 11 lines{
"_id": "ord-4821",
"customerId": "cus-77",
"placedAt": "2026-03-04T10:12:00Z",
"status": "placed",
"shipTo": { "name": "A. Kaur", "city": "Leeds", "postcode": "LS1 4AP" },
"lines": [
{ "sku": "KB-01", "name": "Keyboard", "qty": 2, "unitPrice": 4900 }
],
"total": 9800
}go deeper
Be ready to say that in a document database you decide the shape from the queries the application must serve, and to give one example of a document shaped so a page loads in a single read.
Explain what belongs in an access-pattern list — entry point, returned fields, cardinality, frequency — and walk one operation from requirement to a concrete document shape, including what you left out.
Demonstrate that you price both sides: show a read-optimised shape and state what it costs the write path, then justify the trade with the observed read-to-write ratio and the staleness the product tolerates.
Own the process, not just the outcome: how access patterns get captured and revisited as the product changes, how deliberately-expensive patterns are recorded so nobody silently 'fixes' the hot path, and when a workload shift justifies re-modelling.
## The inversion Most people learn to model by drawing the domain: here are customers, here are orders, here are products, here are the lines between them. Query-first modelling refuses to start there. It starts with a list of the operations the system must perform, and treats the document shape as an *output* of that list. The reason is structural. In a document database the shape you store is the shape you get back. There is no query planner that will reassemble a convenient view out of five stored pieces at low cost — assembling across documents is work the application does, one extra round trip at a time. So the shape is a physical commitment to a particular access profile, and committing before you know the profile is guessing. ## What an access pattern actually is A usable access pattern is more than "we need to get orders". It records: - **The entry point** — what identifier or filter the operation starts from. - **The shape returned** — which fields the caller actually renders or computes with. - **The cardinality** — one document, twenty, or an unbounded list. - **The frequency and the latency budget** — a call on every page load is a different constraint from a nightly report. - **Whether it reads or writes**, and if it writes, what must change indivisibly with it. A typical list for an e-commerce service looks like: *render order detail by order id (very hot, one document, needs lines and totals and the address as captured)*; *list a customer's recent orders by customer id (hot, up to 20, needs id, date, status, total only)*; *find orders by status for the fulfilment queue (moderate, unbounded, needs a summary)*; *recompute lifetime spend per customer (nightly, batch)*. That list, not the entity diagram, decides the documents. Order detail wants lines nested. The customer's recent-orders list wants only summary fields, which means either projecting a few fields out of the order documents or maintaining a compact list — a decision you can only make once you know how hot it is. The nightly batch wants nothing special, because a batch job can afford to scan. ## "Design for the read" does not mean ignore the write The slogan is a corrective to entity-first habit, not a claim that writes do not matter. A shape optimised purely for reading can be miserable to write: if serving the read requires the same value to sit in three documents, every update becomes three writes with a window where they disagree. The honest version of the rule is *design for the dominant access*, and in most interactive systems reads dominate by one or two orders of magnitude, so reads usually win — but you decide that by looking, not by assuming. The useful discipline is to write both columns next to each other: for each candidate shape, what does the hot read cost, and what does the hot write cost? A shape where the hot read is one lookup and the hot write is one document update is the ideal. A shape where the read is one lookup but the write fans out across many documents is a trade you may still take, deliberately, when reads outnumber writes heavily and some staleness is tolerable. ## Working the method 1. **Enumerate operations** from the API surface, the screens, and the jobs. Include the ones that only run at month end; they constrain nothing but they must remain possible. 2. **Rank them** by frequency times latency sensitivity. The top few earn the shape. 3. **Propose a document per aggregate** and check every operation against it: how many accesses does it take, and does it fetch data it discards? 4. **Look for the operations that came out badly** and decide, per operation, whether to accept the extra work, change the boundary, or serve it from a separate derived structure. 5. **Record the losers.** The patterns you consciously made expensive are the ones a future reader will otherwise "fix" by changing the shape and breaking the hot path. ## Failure modes The two symmetrical mistakes are modelling the entity diagram and modelling one screen. Entity-first gives you a tidy set of small documents and an application that issues six reads to render a page. Screen-first gives you a document shaped like today's user interface, which stops fitting the moment the interface changes; the guard is to shape against the *operation* — the data a use case needs — rather than the visual layout it is currently rendered in. The third failure is doing the exercise once. Access patterns drift as features land, and a shape chosen for last year's workload can quietly become the reason a page is slow. Query-first modelling is a habit applied whenever the workload changes materially, not a one-off ceremony at the start of a project. ## How to answer this in an interview State the inversion, give the anatomy of an access pattern, and walk one concrete operation from requirement to document shape. Then volunteer the limit yourself — that reads usually but not always win, and that you check the write cost of every read-optimised shape. Interviewers listen for whether you treat the slogan as a rule or as a trade you have actually priced.
- How is that different from just designing around today's screens?A screen is a rendering; an operation is a data need. Shape against the operation — "the data required to show an order's contents and totals" — rather than the current layout, so a redesign of the page does not invalidate the model. Screens are useful evidence of access patterns, but they are the symptom, not the requirement.
- What do you do with an access pattern that the chosen shape serves badly?Price it. If it is rare or batch, accept the extra reads and note the decision. If it is hot, either move the boundary so it becomes a single access, or serve it from a separate structure maintained alongside the primary aggregate — accepting that this structure now has to be kept current and can lag.
- Does query-first modelling mean you cannot add new queries later?You can, but new queries either fit the existing shape, get supported by an added index, or need a new derived shape. That is precisely why you record which patterns the model deliberately made expensive: a new requirement in that region is a signal to revisit the boundary rather than to bolt on more round trips.
saying these in an interview costs you the question
- Draws the entity diagram first and queries it later
- Treats 'design for the read' as ignore writes entirely
- Lists access patterns without frequency or latency
- Shapes the document to match a screen layout literally
- Assumes the query engine will assemble a convenient view cheaply