skip to content

Reapplying the last known-good revision fixed the bad release but also undid a memory ceiling an operator had raised by hand - why?

level: seniorimportance: should knowfreq 46%

answer

  1. granularity of the move
  2. the apply, not the field
  3. bundled changes travel together
  4. a value in no revision cannot survive
  5. diff the target against live first

basics

~20 s

A revision is the whole spec applied atomically, so reapplying an older one restores every field it holds, including settings that changed after it. Rollback granularity is the apply, never the single field you regret.

solid answer

~50 s

Going back does not mean "undo the last image change". It means "make this whole stored document the desired state again", and that document was written before the ceiling was raised, so it carries the old value for the ceiling, the old copy count and every other field too. There are two ways the raise got lost. If the operator changed the declared spec, the raise became its own revision, and returning to a revision older than that one takes it away with everything else that changed since. If the operator edited the live workload outside the declared source, the raise was never in any revision, and the next apply of any spec overwrites it. Both are cured by the same discipline: one path for changes, and a field-by-field diff of the target revision before you submit it.

code

yaml · 21 lines
yaml
# history, newest last
- revision: 11          # the release before the incident week
  spec:
    contentDigest: "4b81c0a9"
    replicas: 4
    ceiling: { memory: "1GB" }

- revision: 12          # ceiling raised and count grown during Tuesday's incident
  spec:
    contentDigest: "4b81c0a9"
    replicas: 6
    ceiling: { memory: "2GB" }

- revision: 13          # Friday's release - this is the bad one
  spec:
    contentDigest: "e07ad552"
    replicas: 6
    ceiling: { memory: "2GB" }

# going back to 11 also returns replicas to 4 and the ceiling to 1GB
# going back to 12 drops only the bad content reference

go deeper

for a junior

Take away one rule: going back restores the whole stored spec, not just the part that broke. Read what else is in the entry before you ask for it.

for a middle

Explain the granularity: a revision is exactly one apply, so what you bundled together you can only return to together, and a field that never reached the declared source is in no entry at all.

for a senior

Show the practice - diff the target against the live spec field by field, name the differences out loud before submitting, and land every emergency edit back in the declared source the same day.

for a principal

Set the standard: apply granularity is a reliability property, and requiring one change per apply with a single path into the declared spec is what makes rollbacks precise enough to trust under pressure.

## Why the whole spec comes back A rollback on a container platform is not a field-level undo. The stored unit is the whole declared spec as it stood at one accepted apply, and reapplying it makes **all** of that document the desired state again. The platform has no way to express "restore the image reference from revision 11 but keep the ceiling from revision 12" - that mixture is a spec nobody has ever applied, and if you want it you have to author it as a new apply. So the question to ask before any rollback is not "which release am I undoing?" but **"what else changed between that revision and now?"** The answer is everything in the diff, and the diff is the blast radius. ## The two ways the raise disappeared | | The raise was applied to the declared spec | The raise was made outside the declared source | |---|---|---| | Where it lives | its own revision in the history | only in the live workload, in no revision | | What removed it | going back past that revision | the very next apply of any spec | | Is it recoverable from history | yes, the value is in an entry you can read | no, nothing recorded it | | Symptom | the ceiling reverts together with the release | the ceiling reverts on an unrelated change too | Both cases produce the same complaint at 3am and want different fixes. The first is a **granularity** problem: the change you wanted to keep and the change you wanted to drop are separated only by which revision you name. The second is a **provenance** problem: a value that no stored document contains cannot survive a mechanism whose whole job is to make reality match a stored document. ## Keeping the revision boundary equal to the change boundary A revision is exactly one apply, so the precision of your rollbacks is decided when you submit, not when you revert. Three habits do almost all of the work: - **Do not bundle unrelated changes into one apply.** An image change and a ceiling change submitted together can never be separated afterwards. Submitted as two applies, they become two entries and you can return to either boundary. - **Make the declared source the only way a spec changes.** If an emergency edit is applied straight to the live workload, it exists in exactly one place that the next reconciliation is entitled to overwrite. Land the same change in the source as soon as the immediate fire is out - during the incident if you can, in the same hour if you cannot. - **Prefer the newest revision that predates the fault** over the one labelled "last known good". They are often different entries, and the gap between them is the set of changes you are about to lose without meaning to. ## Before you submit the rollback 1. **Diff the target revision against the live spec, field by field.** Not just the content reference - counts, reservations, ceilings, settings, check thresholds. Read the whole diff out loud if somebody else is on the call. 2. **Decide, for each difference, whether you want it.** Anything you want to keep must be re-authored on top of the target, which makes the rollback a new apply built from the old revision rather than a plain reapply of it. 3. **Say what you are about to change** in the incident channel before you press it, in the same words as the diff. Half of the unintended restores in this class are caught by somebody reading that line. ## Why the platform is not wrong to behave this way It is tempting to want a rollback that touches only the field you regret. That mechanism would be worse. A spec is a single consistent statement - a copy count that matches a reservation that matches a ceiling that matches the settings the process expects - and a partial restore produces a combination nobody tested and nobody wrote down. Atomicity per apply is what makes the history meaningful: every entry is a state the system genuinely ran in. The cost of that guarantee is precisely the surprise in this question, and the way you pay it down is by making your applies small enough that the boundary you can return to is the boundary you care about.

  • What should happen to an emergency hand edit once the incident is over?
    It has to be landed in the declared source, so that it becomes part of the desired state rather than a value living only on the live workload. Until that happens it is one apply away from disappearing, and the apply that removes it will usually be something unrelated - which is what makes it hard to diagnose later.
  • You want the old content reference but the current ceiling. How do you get that?
    By authoring it as a new apply: take the target revision, change back the one field you want to keep current, and submit the result. It becomes a new entry at the top of the history, which is honest - that combination is a state the system has not run in before, and the history should say so.
  • Does splitting every change into its own apply have a cost?
    Yes: more entries, more rollouts, and each one is a replacement that takes time and consumes the replacement budget. The trade is precision against churn. Split where the changes are genuinely unrelated, and keep together the fields that only make sense as a set, such as a copy count and the reservation it was sized against.

saying these in an interview costs you the question

  • Says a revert only touches the fields changed in the bad release
  • Believes an edit made outside the declared source survives the next apply
  • Thinks ceilings are outside the spec because the runtime enforces them
  • Bundles unrelated changes into one apply and expects to split them later
  • Submits a rollback without reading what else the target revision holds