skip to content

During a live critical advisory, how do you choose between freezing releases, pinning, and rolling forward?

level: seniorimportance: should knowfreq 49%

answer

  1. three levers, three different jobs
  2. attribution, determinism, completion
  3. how big is the fix diff
  4. who is awake to catch a bad deploy
  5. decide the rule before the incident

basics

~20 s

Freezing stops other changes so the fix is the only thing moving; pinning holds the dependency at an exact known-good version so a rebuild cannot drift; rolling forward ships the fix through the normal train. Choose by exposure and by how much unrelated change the fix drags along.

solid answer

~60 s

They answer different questions. **Freeze** is about attribution and blast radius: if only the fix is moving, a new failure is the fix, and you are not debugging three changes at 7pm. It costs every other team's throughput and the backlog snaps back the moment you lift it. **Pin** is about determinism: hold the affected dependency at an exact fixed version so a rebuild, a cache or a floating range cannot quietly reintroduce the vulnerable one. **Roll forward** is about actually finishing: ship the patched version through the normal pipeline, with all the tests. The decision inputs are exposure, how much unrelated change the patched release carries — a backported patch release is very different from a major bump — and who is around to catch a bad deploy. On a Friday afternoon with a skeleton crew, the usual right answer is mitigate in place now, freeze nothing, and roll forward Monday with real review. What matters most is that this rule was written down before the incident, not argued during it.

go deeper

for a junior

Know what each word means: a freeze stops other changes, a pin fixes an exact version so it cannot drift, and rolling forward ships the patched version through the normal pipeline.

for a middle

Explain what each lever buys and costs — attribution versus throughput, determinism versus stickiness — and why a backported patch release is a much safer roll-forward than a major version bump.

for a senior

Show the combined play: close the path with a mitigation now, scope any freeze narrowly with a lift condition, and separate ending the exposure from finishing the remediation on two different clocks.

for a principal

Own the pre-written rule — what exposure authorises an out-of-hours deploy, who can call and lift a freeze, whether a stage may be skipped — so responders act without convening a committee mid-incident.

## Three levers, three different jobs Under incident pressure these get discussed as if they were alternatives on one axis. They are not; they answer different questions and are frequently used together. ### Freeze Stop merging and releasing anything except the fix. What you buy is **attribution and a narrow blast radius**: when the deploy goes out and something breaks, there is exactly one candidate cause. During an emergency, in which you are skipping soak time and reviewing under pressure, that is worth a great deal. What it costs is everyone else's throughput, and the cost is not linear. A freeze creates a queue, and the queue releases in one lump when the freeze lifts — which is a large, poorly-attributable change landing right after an incident, at the moment your team is most tired. Freezes should therefore be scoped as tightly as the incident: the affected services, not the estate; hours, not days. A freeze with no stated lift condition is a standing tax. ### Pin Declare the affected dependency at an exact fixed version rather than a range, so that nothing — a rebuild, a resolver, a cached layer, a transitive requirement — can move it back. What you buy is **determinism**: the artifact you ship contains the version you decided on, and a rebuild tomorrow produces the same decision. The emergency use is defensive. Under time pressure people rebuild a lot, and a floating range that resolved to the patched version this morning can resolve differently this evening. Pinning removes that class of surprise. The cost is that pins are sticky: an exact version that solved one incident becomes an unreviewed constraint two years later, so a pin needs the same owner-and-expiry treatment as a mitigation. Note also that *how* you force a version that arrives transitively is a separate problem with its own mechanics and its own compatibility risk. ### Roll forward Ship the patched version through the pipeline the normal way. This is the only one of the three that ends the incident, since neither freezing nor pinning removes the flaw by itself — pinning to a *fixed* version does, but pinning as a technique is equally usable to hold the vulnerable version still. The question is what the patched release drags along. A maintainer who publishes a backported patch release hands you a small diff you can read. A maintainer whose fix exists only in the next major version hands you an unrelated set of breaking changes to integrate — during an incident, with the tests you have, at whatever hour it is. That is a materially different risk, and it is often the reason the correct answer is to mitigate now and roll forward properly later. ## The decision inputs 1. **Exposure.** An internet-reachable service processing untrusted input justifies a risky deploy that an internal batch job never does. 2. **Size of the fix diff.** Patch release versus major bump. Read the changelog before you decide, not after the rollback. 3. **Confidence in the train.** If the pipeline has meaningful tests and a staged rollout you can watch, rolling forward is cheap. If the only signal is production, it is not. 4. **Who is awake.** A deploy is only as safe as the people available to notice and reverse it. This is the input that makes Friday afternoon different from Tuesday morning, and it is a legitimate engineering input rather than an excuse. 5. **Whether a mitigation exists.** If you can close the path in minutes, you have converted an emergency deploy into a scheduled one, which is almost always the better trade. ## The Friday-afternoon shape The canonical scenario: a critical advisory lands at 16:00 Friday against a component on a public request path. The tempting answers are both wrong at the extremes — deploying an untested major bump into a weekend with nobody watching, or doing nothing until Monday and leaving a reachable flaw live for 60 hours. The defensible middle: mitigate in place immediately so the path is closed; do not freeze the estate, freeze only the affected services if anything; prepare the version bump and pin it, but do not ship it into an empty weekend unless the exposure genuinely outweighs the deploy risk; roll forward Monday with real review, and verify. The judgment being assessed is whether you can separate *closing the exposure* from *finishing the remediation* and treat them as two decisions with different clocks. ## Why this belongs in a document, not a debate Every input above is arguable, and arguing them at 16:30 on a Friday with an incident channel watching produces the loudest person's answer rather than the best one. Mature teams write a short pre-agreed rule: what exposure level authorises an out-of-hours deploy, who can call a freeze and what lifts it, whether an emergency change may skip a stage and which one, and who signs off on shipping a major version bump as a security fix. The document is not bureaucracy; it is the thing that lets a responder act without convening a committee, and it is the artifact an interviewer is really probing for. ## Rollback, briefly One option people reach for that usually does not apply: rolling *back* to the previous release. For a dependency vulnerability the previous release almost always contains the same or an older vulnerable version, so rollback moves you sideways or backwards on the very thing you are fixing. Rollback is the answer to a bad deploy, not to an advisory.

  • The patched version exists only in a new major release. What changes about your decision?
    The fix now arrives bundled with unrelated breaking changes, so rolling forward means integrating and testing a migration during an incident. That usually flips the answer to mitigate in place now and do the major upgrade as planned work with real review. If exposure is severe enough to force it anyway, ship to a small slice first and keep the mitigation on until the upgrade is verified everywhere.
  • How do you scope and lift a release freeze without stalling the whole organisation?
    Scope it to the affected services rather than the estate, state the lift condition when you declare it — usually the fix verified in production — and name who can lift it. Then expect the queued changes to land in one lump and stage that release deliberately, because a large unattributable batch deploying immediately after an incident is its own risk.
  • Why is rolling back to the previous release rarely the answer to a dependency advisory?
    Because the previous release almost certainly contains the same vulnerable component or an older one, so rollback does not reduce exposure to the flaw. Rollback is the remedy for a bad deploy. The exception is when the vulnerable dependency was introduced by the most recent release itself, in which case reverting genuinely removes it.

saying these in an interview costs you the question

  • Treats freeze, pin and roll-forward as three names for the same action
  • Declares an estate-wide freeze with no stated lift condition
  • Ships an untested major version bump into an unstaffed weekend
  • Thinks pinning removes the vulnerability by itself
  • Proposes rollback when the previous release has the same component

context