A team wants native object-graph snapshots as the cross-service cache format, with per-type version stamps to manage drift; what do you decide?
answer
- detection is not translation
- the cost scales with the graph
- a stamp gets pinned to silence errors
- who owns a contract with no artifact
- short-lived and regenerable is the exception
basics
~20 sRefuse it as a cross-service format. Version stamps only make a mismatch detectable; they cannot translate one, and every type in the graph becomes a versioned contract nobody owns. Keep the family for short-lived, regenerable, single-build use.
solid answer
~50 sThe proposal misreads what a stamp does. A per-type version token turns a silent misread into an explicit failure, which is worth having, but it **translates nothing**: something still has to decide what a reader does with each changed field, and that something is hand-written upgrade code per type per version. The cost scales with the number of types in the graph, not the number of messages, and the graph reaches everything the root points at. Meanwhile the contract is invisible — a refactor in one team's code silently changes what another team can read. I would allow the family only where the bytes are short-lived, regenerable, produced and consumed by one build, and never crossing a trust boundary; for anything that crosses a team, a runtime or a release, I would require an explicit encoding whose contract is an artifact with an owner and a compatibility rule.
go deeper
Take away the rule of thumb: these snapshots are for values that live briefly inside one build, not for anything shared between services or kept for later.
Be able to say what a version stamp actually does — it detects a mismatch, it does not convert old bytes into the new shape — and why that distinction decides the argument.
Show the operational cost: every type reachable from the root joins the contract, each needs stamping discipline and hand-written upgrade handling, and the quickest fix for a failing read is the one that makes it silently wrong.
Decide on ownership and blast radius. An implicit contract spread across every team's type declarations has no owner and no review point, so the boundaries it may cross must be stated as policy rather than left to each pull request.
## What is actually being proposed The proposal is attractive because it looks cheap: no schema to author, no mapping code, no code generation step, and a stamp per type to catch drift. Judging it means separating three questions that get bundled together — what a stamp buys, what the graph costs, and where the boundary should sit. ## What a version stamp buys, and what it does not - **It buys detection.** A stamp or shape token lets a reader notice that the writer's declaration differs from its own, so a mismatch fails loudly instead of producing a half-populated object. - **It does not buy translation.** Knowing that version 3 is not version 5 says nothing about what to do with a field that was split in version 4. Somebody writes that, by hand, per type, per step. - **It is routinely pinned to silence errors.** The stamp is under the type author's control, and the quickest way to make a failing read succeed is to declare the old value. That converts a loud failure into a quiet, half-restored object — the worst outcome of the three. - **It says nothing about intent.** A stamp cannot express that a field became optional, or that an absent value should be treated as a particular default. Those are the evolution rules an explicit contract states and a stamp cannot. ## The cost scales with the graph, not the traffic This is the part teams underestimate. The snapshot covers everything reachable from the root, so every type in that closure is part of the wire contract — including types nobody thinks of as a message, and types owned by other teams. Each one needs a stamp, a discipline about when to bump it, and upgrade handling for each historical shape still present in the cache. The bill arrives in maintenance, long after the decision. ## Which boundaries the family may cross | Boundary | Verdict | Why | |---|---|---| | Inside one process, one build | fine | writer and reader are the same code, alive at the same moment | | Short-lived cache, single service, regenerable on a miss | acceptable with care | a failed read costs a recomputation, not an incident | | Across services or teams | no | the contract has no artifact and no owner, so a refactor breaks a stranger | | Across a release, or into durable storage | no | bytes outlive the code, making the code's shape a permanent contract | | Across a runtime or language | no | reconstruction semantics are private to the writing runtime | | Across a trust boundary | never | the payload names the types the decode path will construct | ## The three questions I would ask before deciding 1. **Can a failed read be turned into a miss?** If the value is regenerable from its source, a version break costs a cold cache. If it is not, a version break costs data, and the family is out. 2. **How long do the bytes live relative to the release cadence?** Bytes whose lifetime is shorter than the interval between deploys never meet code they do not understand. Bytes that persist for weeks certainly will. 3. **Who is allowed to change a type in the graph?** If the answer is any team with a pull request, the contract has no owner and the proposal is already unmanageable. ## The decision, and the thing worth conceding I would refuse a native snapshot as a cross-service cache format and require an explicit encoding for anything crossing a team, a runtime, a trust boundary or a release, because that makes the contract a reviewable artifact with stated compatibility rules. I would not, however, ban the family outright: there is a legitimate use for a short-lived, regenerable checkpoint written and read by one build inside one trust boundary, where the fidelity is genuinely useful and the failure mode is a recomputation. The honest version of the team's argument is not "stamps make it safe" but "we do not want to author a schema" — and that cost is real, one-time, and much smaller than a per-type upgrade matrix maintained forever.
- What would you accept as the legitimate remaining use of this family?A short-lived checkpoint or cache entry written and read by one build inside one trust boundary, regenerable from its source if the read fails. The fidelity is genuinely useful there and the worst outcome is a recomputation.
- Why is pinning a version stamp by hand the dangerous move?It suppresses the only check the scheme has. The read then succeeds against a type whose shape has changed, so values land in the wrong places or silently disappear — a quiet wrong answer instead of a loud failure.
- How do you answer the team's real objection, which is schema authoring cost?Concede it. Authoring an explicit contract is real work, but it is bounded, one-time per message, and reviewable. The alternative is an unbounded upgrade matrix across every type reachable from the root, maintained by whoever inherits the service.
saying these in an interview costs you the question
- Thinks a version stamp lets a reader upgrade old bytes automatically.
- Counts only the root type instead of the whole reachable graph.
- Treats a detectable mismatch as equivalent to compatibility.
- Says the trust-boundary rule is negotiable with enough checks.
- Bans the family outright, ignoring the legitimate short-lived use.