skip to content

An emergency rollback to a release from six weeks ago fails because the delivery system's retention policy already deleted that build's artifact. How would you design artifact retention so this cannot happen?

level: seniorimportance: should knowfreq 44%

answer

  1. age-based expiry applied to everything
  2. classify artifacts before setting a rule
  3. never delete what an environment ran
  4. rollback horizon as an explicit number
  5. rehearse the oldest supported rollback

basics

~20 s

Retain by class, not by age alone: expire per-commit and pull-request builds quickly, exempt any artifact an environment has actually run, and keep released artifacts for at least as long as your stated rollback horizon.

solid answer

~50 s

Age-based expiry applied uniformly is the bug. Split artifacts into classes and give each its own rule: throwaway builds from pull requests and feature branches expire in days; release candidates last weeks; anything that has ever been deployed to a production environment is exempt from age-based cleanup entirely, or retained for a window derived from how far back you have promised you can roll back. That means garbage collection must consult deployment state rather than just timestamps, which in turn means your delivery system has to record which artifact each environment ran and when. Make release identities immutable so a retained artifact cannot be quietly replaced, prune caches and intermediate build outputs aggressively to pay for the retention, and rehearse it — periodically deploy an artifact from the far end of the window to prove the policy holds.

go deeper

for a junior

Know that build outputs are stored with a retention policy and that deleting an old release removes the ability to redeploy it without a rebuild.

for a middle

Explain how to split artifacts into classes with different lifetimes, and why per-pull-request builds and caches are the right things to expire aggressively.

for a senior

Show the operating discipline: reference-aware cleanup driven by deployment history, an explicit rollback horizon, and a rehearsal that proves the oldest supported rollback still works.

for a principal

Own the tradeoff between storage spend and recovery capability across teams, set the horizon as a stated commitment, and align it with migration compatibility and any regulatory retention.

## Why the failure happens Retention is usually configured once, early, as a single age rule — "delete artifacts older than 30 days" — because storage cost is visible and the cost of a missing artifact is not. It stays invisible until the day a regression is traced to a change made two months ago and the only safe move is to put back the last known-good release, which no longer exists as bits. ## Classify before you set a rule One policy across everything is what breaks. Sort the store into classes: - **Throwaway builds** — per-commit, per-pull-request, per-feature-branch outputs. Enormous in volume, valuable for hours or days. Aggressive expiry here is where nearly all the savings are. - **Release candidates** — builds that entered the promotion path. Keep them for the length of the promotion cycle plus a margin. - **Deployed artifacts** — anything an environment has actually run. These are operational assets, not build detritus. - **Released artifacts** — versions handed to customers or to other teams, which may carry contractual or regulatory retention on a completely separate axis. ## The rule that actually prevents the incident *Never expire an artifact that an environment currently runs or ran within the rollback horizon.* Two consequences follow. First, garbage collection must be **reference-aware**: it walks what environments and deployment records point at, not just names and timestamps. A cleanup job that only sees the store will always be able to delete something in use. Second, the rollback horizon must be an explicit number. It falls out of questions you can answer: how long can a regression hide before someone notices; how far back does your data or schema compatibility actually allow you to go; how long is your longest release cycle. Pick the figure, write it down, and set retention to that plus a margin. "Six weeks" stops being a surprise once it is a stated commitment. ## Rebuilding is not a substitute The reflex when the artifact is gone is to rebuild it from its tag. Beyond the fact that a rebuild is not guaranteed to reproduce the tested bytes, it can simply fail during the incident: a dependency version was yanked, a base image tag moved or disappeared, a build credential rotated, the toolchain the project used no longer installs. Treating rebuild as the recovery path means your recovery depends on third parties being available and unchanged, which is precisely the assumption an incident is a bad time to test. ## Making retention affordable Retention arguments are usually really cost arguments, and there are cheaper places to cut: - Expire **build caches and intermediate outputs** on a short clock — they regenerate by design and often outweigh the artifacts themselves. - Cut **per-pull-request builds** to days rather than weeks; nobody rolls back to a PR build. - Rely on **content-addressed deduplication**: successive versions of an artifact frequently share most of their content, so keeping ten releases usually costs far less than ten times one release. - Store older releases in **cheaper, slower storage** if your artifact store supports tiering, accepting a slower first pull in exchange for keeping the bytes. ## Immutability is the other half Retention keeps the artifact; immutability keeps it *the same artifact*. If a release identity can be republished, then "we still have 3.2.0" is not a statement about bytes. Make release identities non-reassignable and have deployments record the content-addressed identity, so the retained artifact and the one that ran are provably the same thing. ## Prove it, do not assume it Retention policy is a backup: untested, it is a belief. Periodically — a quarterly drill is common — take an artifact from the far end of your stated horizon and deploy it to a non-production environment. That single exercise finds all three failure modes at once: the artifact is missing, the artifact exists but no longer starts against current dependencies, or the deployment tooling can no longer express "deploy this specific old identity". ## The adjacent trap Cleanup that is expressed by name rather than by reference can delete something that is still reachable — an untagged object that a retained one depends on. Whatever your store, the safe formulation is the same: collect what nothing references, not what looks old.

  • How do you choose the number that defines your rollback horizon?
    Work backwards from evidence: how long regressions have historically taken to surface, how far back schema and data compatibility genuinely allows you to go, and the length of your longest release cycle. Take the largest, add margin, and publish it as a commitment. It then drives retention, the compatibility window for migrations, and how far back your deployment tooling must be able to target.
  • Why isn't rebuilding the old release from its Git tag an acceptable fallback?
    Two reasons. The rebuild may not reproduce the tested bytes, since dependencies, base images and toolchains resolve at build time. And it may not complete at all during the incident — yanked packages, a moved base image, rotated build credentials. Recovery should not depend on third-party availability at the worst possible moment.
  • What has to be recorded for a garbage collector to know an artifact is still in use?
    A deployment history keyed on the immutable artifact identity: which environment ran which artifact, from when to when. Cleanup then queries that history instead of timestamps alone. Without it the collector can only reason about age and names, which is exactly how an in-use or recently-in-use artifact gets deleted.

saying these in an interview costs you the question

  • We can always rebuild the old release from its tag
  • One age-based expiry rule for every artifact type
  • Storage cost justifies deleting old released builds
  • Only the currently deployed artifact needs keeping
  • Retention is fine because nobody has complained yet

context