skip to content

Under BACKWARD vs FORWARD compatibility, in what order must you upgrade producers and consumers, and why?

level: middleimportance: must knowfreq 70%

answer

  1. BACKWARD → consumers first
  2. FORWARD → producers first
  3. FULL → any order, mixed fleet safe
  4. Mixed fleet = the window that needs the guarantee
  5. Wrong order = deserialization failures mid-rollout

basics

~20 s

BACKWARD: upgrade consumers first, then producers — new consumers can read both old and new data. FORWARD: upgrade producers first, then consumers — old consumers can still read the new data the upgraded producers write.

solid answer

~50 s

The upgrade order falls directly out of the mode's guarantee. BACKWARD guarantees the new schema (reader) can read data written with the old schema (writer), so you deploy the new consumers first; they handle the old records still flowing from not-yet-upgraded producers, and continue working after producers switch. FORWARD guarantees the old schema (reader) can read data written with the new schema, so you deploy the new producers first; the still-old consumers can read the new records, and you upgrade consumers afterward at leisure. FULL satisfies both guarantees, so either order — or a partial, mixed fleet — is safe. The mental shortcut: BACKWARD protects readers, so move readers ahead; FORWARD protects writers' output for old readers, so move writers ahead. Picking the wrong order causes deserialization failures during the rollout window even though each schema is individually valid.

go deeper

for a junior

Recall the two mappings: BACKWARD→consumers first, FORWARD→producers first.

for a middle

Reason through the mixed-fleet window and why the guaranteed direction dictates order.

for a senior

Bring in retention/compaction and transitive variants to argue real-world safety.

for a principal

Tie mode selection to deployment topology, coordination cost, and stream-app changelog reuse across teams.

## Why ordering exists at all During a rollout you have a **mixed fleet**: for a period, some producers and some consumers run the old schema while others run the new one. A topic can simultaneously contain records written by old and new producers, and any of them may be read by old or new consumers. Compatibility mode guarantees *one* direction of reader/writer pairing works; the safe upgrade order is the one that keeps every live pairing inside that guaranteed direction. ## BACKWARD — upgrade consumers first BACKWARD guarantees: **reader = new schema can read writer = old schema**. Rollout: 1. Deploy new consumers. They use the new schema. Producers are still old, so the data on the topic is old-schema. New reader vs old writer → guaranteed OK by BACKWARD. 2. Deploy new producers. Now data is new-schema and consumers are already new. New reader vs new writer → trivially OK. At no point does an old consumer face new data (which BACKWARD does not promise). Hence **consumers first**. ## FORWARD — upgrade producers first FORWARD guarantees: **reader = old schema can read writer = new schema**. Rollout: 1. Deploy new producers. Data becomes new-schema. Consumers are still old. Old reader vs new writer → guaranteed OK by FORWARD. 2. Deploy new consumers. New reader vs new writer → OK. At no point does a new consumer face old data unless the new schema can also read it — FORWARD does not promise that, so you must finish the producer rollout's data being new before... actually new consumers reading residual old data is the gap, which is why FORWARD upgrades **producers first** and you accept that old data lingering needs care. Hence **producers first**. ## FULL — any order FULL = BACKWARD AND FORWARD. Every reader/writer pairing in both directions is guaranteed, so a mixed fleet in any order is safe. This is why teams that cannot coordinate deploy timing across many services often standardize on FULL (or FULL_TRANSITIVE). ## Edge cases - **Residual old data**: with long retention or compacted topics, old-schema records persist. A non-transitive mode only guarantees the immediately previous version, so a third schema version might not read the first. Use the **_TRANSITIVE** variant to guarantee across the whole history. - **Consumer groups resetting offsets** to the start re-expose old data, making transitive guarantees matter even under BACKWARD. - **Stateful stream apps** (Kafka Streams) reading their own changelog topics effectively are both producer and consumer of evolving schemas; FULL/FULL_TRANSITIVE is the typical safe choice. ## The memory shortcut BACKWARD → **B**ack → readers move first (consumers first). FORWARD → **F**orward → writers move first (producers first). FULL → free order.

  • Your team can't coordinate deploy timing between many producer and consumer services. Which mode reduces ordering risk?
    FULL (or FULL_TRANSITIVE), because it guarantees both directions, so any deploy order and a mixed fleet are safe.
  • Under BACKWARD, what breaks if you upgrade producers first by mistake?
    The still-old consumers face new-schema data, which BACKWARD does not guarantee they can read, causing deserialization errors until consumers catch up.

saying these in an interview costs you the question

  • Saying BACKWARD means producers first — it's consumers first.
  • Claiming the order doesn't matter for any mode except FULL — it matters for BACKWARD and FORWARD.
  • Ignoring residual old data / retention when arguing an upgrade is safe.
  • Assuming a single global mode covers all subjects regardless of topology.

context