skip to content

Which compatibility guarantee would you require for a long-retention event stream versus a short-lived internal channel, and why?

level: principalimportance: should knowfreq 38%

answer

  1. guarantees are bought, not free
  2. retention sets the reach
  3. deploy control sets the direction
  4. the strongest rule freezes the schema
  5. tightening never binds past records

basics

~20 s

Match the guarantee to how long data and stragglers outlive a release. A long-retention stream needs a transitive guarantee so current readers can replay all of history; a short-lived channel drained between releases can run honestly on the weaker per-step form.

solid answer

~50 s

Treat the guarantee as something you **buy**, not a default. It costs schema authors design freedom permanently, and it buys the removal of cross-team deployment coordination. Two axes set the price. **Reach** comes from retention: if records or stragglers outlive a release, you need the transitive form; if the channel drains between releases, the per-step form describes the pairings that actually occur. **Direction** comes from control: if you can reliably order one side first, require the direction that covers the resulting window; if you cannot — independent consumers, a two-way exchange, changes that must be rollable — require both. So a long-retention stream replayed from its start lands on a transitive guarantee, usually backward; a short-lived internal channel between services one team deploys together needs far less. And a rule adopted today binds future versions only — it never brings already-written records into line.

go deeper

for a junior

Understand that different channels can carry different compatibility rules, and that the rule is chosen deliberately rather than being a property the schema happens to have.

for a middle

Be able to say what each guarantee requires and to name one concrete situation — replay of a retained log, a fleet you cannot sequence — where the weaker one would not be enough.

for a senior

Derive the requirement from the two facts that decide it: how far back a reader can be asked to reach, and whose upgrade order you can actually enforce. Then say what the choice costs.

for a principal

Own the policy and its price: which channels deserve which strength, what design freedom you are taking from schema authors in exchange, and the fact that a rule adopted today governs only what is published from today.

## The guarantee is a purchase, not a default The reflex answer — "require the strongest guarantee everywhere, it is the safe choice" — is the one an interviewer is probing for. A compatibility guarantee is not free. It is a **standing constraint on every future change to that schema**, levied on the people who author it, and it is paid in order to buy something specific: the right to stop coordinating deployments across teams. Under both directions held over the whole history, an author effectively cannot remove or repurpose anything an older reader might still rely on, and every addition has to be something older readers can live without. The schema accretes and nothing ever leaves. That is an acceptable price on a channel where the coordination it removes is genuinely expensive, and a pure loss on a channel where nobody was going to be out of step anyway. ## Two axes, decided separately The decision is not one choice from a menu of three. It is two independent questions: 1. **Reach — how far back must a reader be able to go?** This is a retention question, not a schema question. Answer it with: how long are records kept, does anything replay or backfill from the start, and how far behind can the slowest consumer be? If the answer to any of those spans more than one release, the claim has to quantify over the history, not the previous version. 2. **Direction — whose upgrade order can you actually enforce?** If you can guarantee readers move first, backward covers you. If writers move first, forward does. If you cannot enforce either — many independent consumers, a two-way exchange, or a requirement that changes be rollable — you need both. ## Applying it to the two channels | | Long-retention event stream | Short-lived internal channel | |---|---|---| | Oldest data a current reader may see | many versions old | at most the current one | | Who reads it | many teams, some far behind | one team's own services | | Enforceable deploy order | no | usually yes | | Sensible requirement | transitive, direction chosen from who can be ordered; both if nobody can | per-step, one direction, or a documented cutover | On the stream, the deciding fact is that replay points today's readers at years of history, so the plain per-step check simply does not describe the decodes that happen. On the internal channel, requiring the strongest guarantee constrains every future change to buy protection against a window that is drained away in minutes. ## The failure modes on both sides - **Over-specifying** looks free and is not. It shows up years later as a schema nobody can clean up, full of fields kept alive for readers that were decommissioned long ago, and as teams routing around the rule with a parallel channel because the rule made the honest change impossible. - **Under-specifying** shows up as an outage during an ordinary deploy, or worse during a rollback, when a pairing nobody checked finally occurs. - **Specifying without a range** is the quiet one: a rule with no stated reach is enforced against whatever the checker happens to compare, and nobody notices the gap until a replay. ## What a rule adopted today can and cannot do This is the part that separates a considered answer from a confident one. Tightening a channel's requirement constrains **every version published from now on**. It does **not** make already-written records comply: they were produced under the old rule and are unchanged by the decision. If history genuinely has to satisfy the stronger claim, the honest options are to re-encode it, or to state a bound — readers are not required to reach past some version — and to make that bound real by retiring the data behind it. The same asymmetry applies to loosening a rule: relaxing it going forward does not un-promise what consumers already built against, and consumers that assumed the stronger guarantee will keep assuming it until told otherwise. ## How to present the decision A lead is expected to give an answer with a shape, not a slogan: 1. Classify each channel by retention and by whether an upgrade order is enforceable. 2. Derive reach and direction from those two facts rather than from a house default. 3. Write the resulting requirement down **with its version range**, because an unwritten range is one nobody maintains. 4. Revisit it when either fact changes — retention is extended, a second organisation starts consuming, the channel becomes two-way. How such a rule is enforced at publication time is a separate matter with its own mechanics. What a lead owns is which rule each channel deserves, and being able to say what it costs.

  • What does requiring the strongest guarantee cost the people who author the schema?
    Design freedom, permanently. With both directions held over the whole history, an author can effectively never remove or repurpose anything an older reader may rely on, and every addition must be something older readers can do without. The schema only accretes, and cleaning it up stops being possible.
  • Can you tighten a channel's requirement after the fact?
    Going forward yes, retroactively no. A stricter rule constrains the next version onward, but records already written were produced under the old rule and may already violate the stronger claim. If history must comply, the options are re-encoding it or declaring a bound on how far back readers are required to reach.
  • How do you tell when a weaker guarantee is the right answer rather than laziness?
    Check whether the window the guarantee exists to cover actually occurs. If a channel retains nothing across a release and has no reader on an older version, the stronger rule protects against a pairing that never happens while still taxing every future change. Say that out loud, and write down the retention assumption it rests on.

saying these in an interview costs you the question

  • Requires the strongest guarantee everywhere and calls it a safe default
  • Picks a direction without asking whose deploy order is enforceable
  • Assumes tightening the rule fixes records already written
  • Ignores retention when choosing between per-step and transitive
  • Treats the guarantee as free rather than a cost on schema authors