skip to content

When publishing one node of a deeply linked build graph to another team, how deep should the serialized form reach?

level: principalimportance: should knowfreq 34%

answer

  1. serialize this node is underspecified
  2. the closure is most of the graph
  3. consistency against coupling
  4. identifier plus a display field
  5. cut at the unit of work

basics

~20 s

Cut at the boundary of what the consumer cannot act without. Embed that subgraph as one consistent snapshot; emit stable identifiers for everything beyond it, plus the few fields needed to avoid an immediate follow-up fetch, and mark which snapshot the references were consistent with.

solid answer

~50 s

Deep and shallow are the two ends, and both fail at scale. A deep cut embeds the transitive closure: one consistent snapshot, no follow-up fetches, but a payload that grows with the graph and a contract that now includes every reachable type, so any of their schema changes becomes your compatibility problem — and it may carry nodes this consumer is not entitled to see. A shallow cut emits identifiers only: small, stable, minimally coupled, but it pushes fetch amplification onto the consumer and its reads can tear across versions. The defensible middle is to cut at the aggregate boundary: embed what the consumer cannot act without, reference the rest by a stable identifier plus a couple of denormalised fields for display, and stamp the message with the point in time the embedded snapshot was consistent at. Which fields those are comes from the consumer's use cases, so this is a contract conversation, not an encoder setting.

go deeper

for a junior

Recall that following every reference out of one node can reach almost the whole graph, so publishing a node means deciding where to stop rather than sending everything it touches.

for a middle

Contrast embedding the reachable subgraph with emitting identifiers, and name the concrete costs on each side: payload size and coupling against extra fetches and inconsistent reads.

for a senior

Diagnose a cut from its symptoms — consumers always fetching the same references, or your reviews firing on types you do not own — and describe the reference shape that removes most of the amplification.

for a principal

Own the trade explicitly: which consumers get which depth, what consistency the message promises, and that a deep cut makes another team's schema changes your compatibility obligation.

## The question behind the question A node in a build graph reaches, transitively, most of the graph. 'Serialize this node' is therefore not well defined until someone says where the walk stops. That decision is the single most consequential one in a published message, because it fixes the payload's size, its consistency guarantee, and the set of schemas you have coupled yourself to. ## The two ends **Deep cut — embed the transitive closure.** - The consumer gets everything in one read, with no follow-up calls and no partial view. - The embedded subgraph is a **consistent snapshot**: every node was read at one moment, so the consumer never sees a mixture of versions. - But the payload grows with the graph, and the growth is usually worse than linear, because a shared node reached by several paths is embedded once per path unless identity indirection is in the contract. - Every reachable type's schema is now part of **your** published contract. When a team three hops away changes their node's shape, your consumers feel it, and you are the one who fields it. - You may embed nodes this consumer is not entitled to read, turning a shape decision into an authorization decision. **Shallow cut — emit identifiers.** - Small, stable payload; your contract covers your node and nothing else. - The consumer resolves what it needs and caches each referenced node independently, which is efficient when consumers need very different slices. - But a consumer that needs the neighbourhood now issues one fetch per reference, and a list of such nodes multiplies it. - Its reads **tear**: each referenced node is fetched at its own moment, so the assembled view may mix versions that never coexisted, and nothing in the payload says so. | Dimension | Deep cut | Shallow cut | |---|---|---| | Payload size | Grows with the reachable graph | Bounded by your node | | Consistency | One snapshot across the subgraph | Per-reference, may tear | | Schema coupling | Every reachable type joins your contract | Your node's type only | | Consumer round trips | One | One per reference it resolves | | Caching | All or nothing, per message | Per referenced node, reused | | Authorization | Must be decided per embedded node | Enforced where each node is fetched | ## The middle that usually wins 1. **Find the boundary of the consumer's unit of work.** Embed the nodes it cannot act without — the ones whose absence makes the message useless — and reference the rest. Where both sides already share a notion of an aggregate, that boundary is the obvious cut. 2. **Make each reference useful on its own.** A bare identifier guarantees a follow-up fetch. A stable identifier plus the two or three fields the consumer needs to display or filter — a name, a kind, a status — removes most of the amplification while keeping the other team's schema out of your contract. 3. **Say what the snapshot means.** Stamp the message with the moment the embedded subgraph was read, so a consumer that resolves a reference later can tell that what it fetched may be newer than what it was given. 4. **Pick the cut per consumer type, not per encoder.** Several messages carrying different depths of the same graph is an honest design; one message that serves everybody tends to be deep for those who wanted small and small for those who wanted deep. When the same graph must serve genuinely different consumers, the alternative to guessing is to let the consumer state the depth it wants and to answer only within a published maximum. That trades a simpler payload for a more complex contract, and is worth it only when the consumers' needs really do diverge. ## Signals that the cut is wrong - Consumers routinely fetch the same set of references immediately after receiving your message — the cut is too shallow for their unit of work. - Your compatibility reviews keep being triggered by changes to types you do not own — the cut is too deep. - Payload size tracks the graph's growth rather than the entity's — the closure is being embedded where references belong. - Two consumers disagree about the same entity because each resolved its references at a different moment — consistency was never stated. ## What an interviewer is listening for That you refuse the framing of one universally right depth; that you name the three axes that actually decide it — consistency, coupling and amplification — and pick a cut you can defend for a stated consumer; and that you recognise the deep cut's quiet cost, which is that you have adopted other teams' schema changes as your own compatibility obligation.

  • What does a deep cut cost you that does not show up in the payload size?
    Compatibility obligations you do not control. Embedding a type puts its shape in your published contract, so a change by the team that owns it becomes a breaking change for your consumers and a review you have to run. It also forces an authorization decision per embedded node, since the consumer now receives data it was never checked against.
  • How would you keep a shallow cut from causing a fetch per reference?
    Carry enough denormalised context in each reference for the common case — a stable identifier plus the name, kind or status the consumer displays — so a fetch is needed only when it truly acts on the referenced node. Batch resolution helps too, but the real lever is knowing which fields the consumer needs, which makes this a contract conversation.
  • A consumer complains that two of its views of the same entity disagree. What do you suspect?
    References resolved at different moments. A shallow cut gives no consistency guarantee across its references, so two assembled views can mix versions that never coexisted. Stamping the message with the moment the embedded part was read makes the tear visible, and moving the nodes that must agree inside the embedded cut removes it.

saying these in an interview costs you the question

  • Claims one depth is correct for every consumer
  • Embeds the transitive closure and calls it simpler
  • Ignores that embedding adopts other teams' schema changes
  • Emits bare identifiers and calls the amplification the consumer's problem
  • Assumes references resolved separately give a consistent view
  • Treats the cut as an encoder setting rather than a contract