A payload version travels on a queue and into archived files, where no request-and-response exchange exists. Which versioning strategy still works, and why?
answer
- nobody to negotiate with here
- the marker rides with the bytes
- the reader dispatches on the version
- retention sets the lifetime, not deploys
- replay reads the whole retention span
basics
~20 sOnly a self-carried marker works: the version must sit in the bytes, because negotiation needs a counterpart to ask and an address to ask it at, and a consumer reading a queued message or an archived file has neither.
solid answer
~50 sMedia-type negotiation and address-based versioning both assume a synchronous exchange: someone states a preference and someone answers it. A queue consumer receives whatever was published, and a file reader opens whatever was written years ago, so neither has anyone to negotiate with. The version therefore has to be **self-carried** - a small field or schema id at a known position in the message - and dispatch moves into the reader, which must handle every version still present. Two consequences follow. The lifetime of a version is set by **data retention**, not by client upgrades: a record in an archive outlives every running caller. And a single producer fans out to many consumers at different versions at once, so "run the old version side by side" would mean duplicating the pipeline per version rather than adding a route.
code
json · 10 lines{
"v": 2,
"type": "order.placed",
"occurredAt": "2026-09-18T10:31:05Z",
"data": {
"orderId": "a41f",
"amountMinor": 1299,
"currency": "EUR"
}
}go deeper
Recall that a message on a queue or a record in a file has to say inside itself which version it is, because there is nobody for the reader to ask.
Explain why negotiation needs a request and a counterpart, and what a reader does instead: find the marker first, dispatch on it, and fail loudly on one it does not know.
Show that retention, not deployment, sets how long a version must stay readable, and that a replay makes every version in the window live again today.
Decide the policy: how long readers carry old paths, whether archives are migrated or read in place, and why fan-out makes writer compatibility cheaper than per-version streams.
## Why negotiation is unavailable here Negotiating a representation requires two things: a counterpart who is present when you ask, and a request you can attach the preference to. A synchronous call has both. Two common channels have neither: - **A queued or published message.** The producer wrote it and moved on, possibly before the consumer existed. There is no request to carry an `Accept` preference, and no producer waiting to honour one. - **Data at rest.** An archived file is read minutes or years after it was written, quite possibly by a program that did not exist at write time. Whatever the synchronous face of the same system does, these channels force the version to travel **with the bytes**. ## What self-carried means in practice - A small marker at a known position - a version number or a schema identifier - that a reader can find without first knowing the shape. - **Dispatch in the reader.** The consumer inspects the marker and selects an interpretation, which means the consumer contains code for every version it may legitimately meet. - **An explicit rule for unknown versions.** A newer marker than the reader understands must be a loud, routed failure - park the message, alert - never a best-effort guess, because guessing turns a rollout bug into silent corruption that is discovered downstream. ## The consequences that catch people out | Question | Synchronous channel | Queue or archive | |---|---|---| | Who selects the version? | The caller, per request | The producer, once, at write time | | How long must an old version be readable? | Until the last caller upgrades | Until the last **record** ages out of retention | | What does "run both versions" mean? | Another route on one service | Another stream, or a reader that handles both | | Who feels a bad rollout first? | The caller, immediately | Whoever replays the backlog, later | **Retention, not deployment, sets the lifetime.** This is the point most candidates miss. You can upgrade every producer and every consumer on Monday and still have seven years of records in the archive whose interpretation depends on a version your code no longer contains. Either the reader keeps the old paths for as long as the data is retained, or the data itself is migrated - and migrating an archive is a batch rewrite with its own verification, not a deployment. **Fan-out makes side-by-side expensive.** One producer serves many consumers, each upgrading on its own schedule. Publishing every version to its own stream multiplies the pipeline - storage, ordering, lag monitoring, replay - so the usual answer is one stream where the writer stays compatible and the readers dispatch on the marker. That pushes the work onto compatibility discipline rather than onto routing. **Replay is the sharp edge.** Re-processing a backlog means today's consumer reading bytes written across the whole retention span, so every version in that window has to be decodable *now*. A team that removed an old reading path the week after the last old producer stopped writing will discover this during an incident replay, which is the worst possible moment. ## What to do about it 1. Put a version or schema marker in every message and every archived record from the first release, before you need it. Retro-fitting a marker onto unmarked data means guessing from shape, which is exactly the circularity the marker exists to avoid. 2. Keep readers able to decode everything within the retention window, and tie removal of a reading path to a retention date rather than to a deployment. 3. Prefer compatible writer changes here even more strongly than on a synchronous path, because you cannot ask consumers what they want - you can only find out what they could not handle. 4. Route unknown versions to a parking area with an alert, and treat anything landing there as a rollout defect rather than as data. Where ecosystems differ is only in what the marker looks like - an inline field, a fixed-width prefix, or an identifier resolved elsewhere - and not in the constraint, which comes from the absence of a counterpart rather than from any format.
- When can a consumer safely delete the code that reads an older version?When no record at that version remains readable anywhere it may be asked to read - live stream, backlog and any archive within its retention period - not when the last producer stopped emitting it. Tie the deletion to a retention date, and confirm with a counter on the old decoding path.
- Why is publishing each version to its own stream usually the wrong answer to fan-out?Because it duplicates the pipeline rather than a route: storage, ordering guarantees, lag monitoring, replay tooling and consumer-group management all fork per version. One stream with a marker and version-dispatching readers keeps the operational surface single and pushes the effort into writer compatibility, which is cheaper to hold.
- What should a consumer do with a message whose version marker is newer than anything it knows?Stop, park it and alert. Processing it on a best-effort basis produces plausible but wrong results that are discovered far downstream, while dropping it loses data silently. A parked message with an alert keeps the record and turns the condition into a rollout defect someone can fix.
saying these in an interview costs you the question
- Assumes a consumer can negotiate the version it prefers
- Ties removal of an old reading path to the producer's upgrade
- Forgets that a replay reads the whole retention window
- Processes an unknown newer version on a best-effort basis
- Duplicates the whole pipeline per version to avoid dispatch
- Leaves messages unmarked and infers the version from shape