In an underwriting platform, what makes a retired model version un-rebuildable while its artifact still sits in object storage?
answer
- servable is not rebuildable
- the reference graph, not the artifact
- snapshot retention set by storage cost
- retirement is not expiry
- exercise the window on a schedule
basics
~20 sRebuildability depends on the pins outliving the artifact. Most often the training-data snapshot was deleted under a shorter retention policy; a withdrawn base image, an unavailable dependency version or a rewritten feature definition do the same.
solid answer
~40 sStoring the artifact makes a version **servable**; storing its pins and everything they point at makes it **rebuildable**, and the second decays on a different clock. Four things expire independently: the **training-data snapshot**, usually deleted under a retention policy set for data volume rather than for audit; the **base image**, withdrawn or garbage-collected from its store; a **resolved dependency version** that is no longer fetchable; and the **feature-definition version** the training job used, if definitions are edited in place rather than versioned. The fix is to make the rebuildable window an explicit commitment - a stated period during which every pin of a promoted version is retained together - and to verify it by rebuilding sampled versions on a schedule, so decay is discovered by a cron job rather than by a dispute.
go deeper
Remember that keeping the trained artifact is not the same as keeping the ability to rebuild it; the inputs it was built from expire on their own schedules.
Name the nodes in the reference graph that can disappear - the data snapshot, the base image, a dependency version, a feature definition - and why each breaks a rebuild.
Separate retirement from expiry, verify the window by rebuilding sampled versions on a schedule, and expire a version's pins together so no record claims a guarantee it cannot meet.
Set the window from the dispute exposure the business actually carries, then make it affordable by tiering versions and deduplicating shared inputs by content digest.
## Rebuildable is not the same as stored A retired model version can be in three distinct states, and teams routinely conflate the second and third: | State | What the platform holds | What it can answer | |---|---|---| | Described | The lineage entry and the pin set | What the version was built from | | Servable | The artifact itself | What the version outputs today | | Rebuildable | The pins **and** everything they reference | Whether re-running those inputs reproduces it | An artifact sitting safely in object storage delivers the middle row only. Rebuildability is a property of the whole reference graph, and it fails the moment any node in that graph disappears - regardless of how carefully the artifact itself was preserved. ## The four things that expire 1. **The training-data snapshot.** The most common cause. Snapshots are large, so their retention is set by whoever owns the storage bill, on a clock that has nothing to do with how long a decision can be disputed. A ninety-day data retention and a seven-year audit obligation cannot both be right. 2. **The base image.** Images get garbage-collected, or a tag is cleaned up, and the digest the pin set records no longer resolves to anything fetchable. 3. **A resolved dependency version.** A specific version may be withdrawn from the registry it was fetched from, or the mirror that held it is retired. 4. **The feature-definition version.** If feature definitions are edited in place rather than versioned, the training job's transformations cannot be reconstructed even with perfect data and code, because the definition that produced the columns no longer exists in the form it had. A fifth, quieter failure: the pins are all intact but the **compute** they assume is gone - an accelerator generation that has been decommissioned. The rebuild can still run, but no longer in an environment that could produce bit-identity. ## Staged promotion sets the clock Platforms differ in how many lifecycle stages they name and what they call them, but the mechanism is the same everywhere: a version moves from a state where it is a candidate, to a state where it takes live traffic, to a retired state, and each transition changes the platform's obligations toward it. - **Promotion to live** is what starts the rebuildability clock, because from that moment the version can produce decisions someone may dispute. - **Retirement** stops the version taking traffic. It is explicitly *not* the end of the obligation - a version retired today may still have issued a decision last week whose dispute window runs for years. - **Expiry of the rebuildable window** is a third, separate event, and it is the only one that may delete pins. Because it is the destructive step, it should be the one that requires a deliberate decision rather than a default policy on a storage bucket. The common failure is having only two events instead of three: the version is retired, and "retired" quietly becomes permission to clean up everything attached to it. ## Choosing the window The window is a business input, not a storage decision. It is bounded below by however long a decision made by that version can be challenged, plus the time needed to answer a challenge. Two refinements make it affordable: - **Tier it.** Not every version needs the same window. Versions that took production traffic get the full commitment; candidates that never left evaluation need far less. - **Deduplicate the heavy nodes.** Many versions train from the same snapshot; retaining that snapshot once by content digest serves all of them and makes the cost sublinear in the number of versions retained. ## Retiring a version honestly When the window does expire, retire the pins **together** and record that you did. A pin set pointing at a deleted snapshot is worse than no pin set: it looks like a guarantee and is not one. Mark the version as described-but-not-rebuildable, keep the small lineage entry - it is cheap and it still lets an audit say precisely what the version was built from - and let the answer to a very old dispute be a precise description rather than a failed rebuild attempt discovered halfway through. ## Proving the window rather than asserting it Rebuildability decays silently: nothing errors on the day the snapshot is deleted. The only way to know the commitment still holds is to exercise it - pick a live or recently retired version on a schedule, rebuild it from its pins, and alert when any referenced node is missing. That turns a policy statement into a monitored property, and it moves the discovery of every expired pin from an audit conversation to a routine alert.
- Why is retaining a pin set that points at a deleted snapshot worse than deleting the pin set too?Because it misrepresents the guarantee. The record looks complete, so the platform reports the version as rebuildable and nobody revisits it until someone tries - typically under dispute, with a deadline. Expiring the pins together and marking the version described-but-not-rebuildable states the true position up front, which is a weaker claim but an accurate one that the team can plan around.
- How do you keep the rebuildable window affordable when many versions share one training-data snapshot?Retain the heavy nodes by content digest and reference them from every pin set that used them. Twenty versions trained on the same snapshot then cost one copy, not twenty, and the retention cost grows with distinct inputs rather than with version count. Tiering helps too: versions that took production traffic get the full window, while candidates that never left evaluation get a much shorter one.
saying these in an interview costs you the question
- Treats a stored artifact as proof the version can be rebuilt
- Lets data retention policy silently set the audit window
- Deletes everything attached to a version the moment it is retired
- Edits feature definitions in place rather than versioning them
- Asserts a rebuildable window that has never once been exercised