skip to content

You roll a service back to its previous version an hour after a bad release. During that hour the new version wrote records containing a field the older version knows nothing about. What problems does the rollback now create, and what limits how far back you can safely go at all?

level: seniorimportance: should knowfreq 40%

answer

  1. the code went back, the data did not
  2. unknown field silently dropped on rewrite
  3. poisoned cache and unreadable messages
  4. tolerant reader in both directions
  5. horizon is the oldest still-compatible version

basics

~20 s

Data written by the newer version can be unreadable, silently dropped, or overwritten by the restored older code. Your real rollback horizon is the oldest version still compatible with today's schema, message formats and cached data - usually one release back.

solid answer

~50 s

Rolling back the code does not roll back the world it changed. Rows, events, cached objects and queued messages written during that hour carry the new shape, and the restored version meets them cold: strict deserializers throw, an unknown enum value hits a default branch, and a read-modify-write cycle can silently drop the new field on the next update - a quiet data loss that is far worse than a visible error. Poisoned caches and un-parseable messages keep failing after the rollback until you version the cache keys or park the bad messages. So the honest answer to "how far back can we go?" is not "any release" - it is the oldest version still compatible with today's schema, message formats, external contracts and credentials, which for most services is one release back and occasionally a handful. Make the old version a tolerant reader, and know your horizon before you need it.

go deeper

for a junior

Understand that a rollback restores the code only, and that records the newer version wrote are still there afterwards for the old version to deal with.

for a middle

Explain the concrete failure modes - strict deserialization, unknown enum values, silent field loss on read-modify-write - and how tolerant readers and additive schemas prevent them.

for a senior

Demonstrate that you know your service's rollback horizon and what shortens it, that you plan for poisoned caches and dead-lettered messages, and that you treat post-rollback data repair as part of the incident.

for a principal

Set the compatibility contract across services - additive-only during the rollback window, unknown-field preservation, versioned cache and index keys - and decide how much dual-compatibility cost each service is required to carry.

## The rollback moves the code, not the data A rollback restores a binary. It does not restore the database rows, the topic offsets, the cache entries or the third-party side effects the new version produced while it was live. Those are now inputs to the old version, and the old version was written before they existed. The interesting failures are the quiet ones. ## Four ways forward-written data hurts **1. Hard read failures.** A strict deserializer that rejects unknown fields will throw on every record the new version wrote. Errors cluster on recently touched entities, which makes the symptom look random until you notice they all share a write timestamp inside the incident window. **2. Silent field loss.** This is the dangerous one. The old code reads a record, ignores the field it does not know, and writes the whole object back. The new field is gone - no error, no alert, and no way to tell later which rows lost it. If the new field carried anything customers can see, you now have a data-integrity incident hiding behind a resolved availability incident. **3. Unknown enumerated values.** The new version introduced a status, type or state that the old code has no branch for. Depending on the language it either throws or falls into a default that means something different, quietly mis-classifying those records. **4. Residue that keeps failing after the rollback.** Caches hold serialized objects in the new layout; queues hold messages in the new schema; a search index holds documents with a new mapping. The rolled-back code keeps hitting them and keeps failing, so the SLI does not recover and it looks like the rollback did not work. The fixes are unglamorous: version cache keys so a release change invalidates them implicitly, park un-parseable messages in a dead-letter destination rather than blocking the consumer, and be prepared to drain or reprocess. ## Designing so the old version survives the new one's leftovers The property you want is **compatibility in both directions**, and it is engineered before the release, not after: - **Tolerant reader.** Both versions ignore fields they do not recognise instead of failing on them. This is what makes additive change safe. - **Preserve unknown fields on write.** If your serialization can round-trip unrecognised fields, the silent-loss problem disappears. If it cannot, avoid read-modify-write over whole objects or accept that rollback loses those fields. - **Additive-only message and event schemas** during the window in which you might roll back, with a registry check if you have one. - **Version keys, not just values.** Cache keys and index names that include a schema or release version make stale-format entries unreachable rather than poisonous. - **Never make a new field required immediately.** Optional first, then required after the rollback window. ## What your rollback horizon actually is Teams say "we can always roll back". Almost nobody can. The horizon is the oldest version that can still run correctly against *today's* state, and it is shortened by: - **Completed contract steps** - once a column, field or endpoint the older code used is gone, that code cannot run. - **New required fields or constraints** the older version does not populate. - **Message and API contracts** - a downstream dependency that has moved on, or a partner API version the old client no longer speaks. - **Rotated or expired credentials and pinned certificates** baked into the older artifact. - **Artifact availability** - a release whose build no longer exists is not a rollback target. For most services the honest horizon is one release back, sometimes a small number of releases, and it changes every time a contract step lands. The practice worth naming in an interview is: state the horizon explicitly, keep it in the release notes or the service's operational metadata, and verify it - the cheapest verification is that your pre-production environment routinely runs the previous version against the current schema. ## Cleaning up afterwards A rollback that lost or mangled data is not finished when the SLI recovers. The follow-up work is part of the incident: identify the affected write window, quantify the rows or messages involved, decide whether to repair them from another source (an event log, an audit table, a backup) or accept the loss, and record it. Teams that skip this step discover it months later as a support ticket nobody can explain.

  • Why is silent field loss during a rollback considered worse than an outright deserialization error?
    An error is loud, bounded and alerting - you see it immediately and can stop it. Silent loss produces no signal: the old code reads a record, drops the field it does not recognise, and writes it back. The damage spreads with normal traffic, no dashboard moves, and by the time anyone notices you may not be able to identify which records were affected or reconstruct what they held.
  • After the rollback, errors persist only for requests that hit the cache. What is happening, and how do you avoid it next time?
    The cache still holds objects serialized in the new version's format and the key does not encode the format, so the restored code keeps reading entries it cannot parse. Immediately, you flush or bypass the affected keyspace. Structurally, include a schema or release version in the cache key so a version change makes stale entries unreachable rather than poisonous - invalidation becomes implicit instead of a manual step under pressure.
  • How would you make your team's actual rollback horizon a known number rather than an assumption?
    Derive it from the constraints: the last completed contract step, any new required fields, changed external contracts, and credential or certificate lifetimes in older artifacts. Record it per service alongside the release, and verify it by regularly running the previous version against current pre-production state. The point is that the number changes with every migration, so it needs an owner rather than a one-off answer.

saying these in an interview costs you the question

  • Rolling back the code also rolls back the data it wrote
  • Unknown fields are always ignored safely by any deserializer
  • We can roll back to any release we still have an artifact for
  • The rollback failed because the errors did not stop immediately
  • Once the SLI recovers, the incident's data effects are finished

context