A rerun overwrote the stored artifact of an already-promoted claim-severity model in place - what breaks?
answer
- a version names bytes, not a location
- write-once, keyed by content digest
- metrics and approval describe the old bytes
- fleet splits as instances recycle
- verify the digest on load, not only on write
basics
~20 sEvery claim about that promoted version becomes false. Its recorded metrics and its approval now describe bytes that no longer exist, the version you would roll back to may be the one overwritten, and two serving instances loading the same version can run different models with no version change to explain it.
solid answer
~50 sA model version has to be a **write-once, content-addressed** artifact: the registry stores a digest, not a mutable path. Overwriting in place breaks four things at once. The evaluation metric recorded on the version no longer describes the bytes being served. The approval that promoted it signed off on something else. The artifact you would fall back to may be exactly the one that was replaced. And a serving instance that loads at start-up gets a different model from one that loaded yesterday, so predicted claim severity shifts with **no version change anywhere to explain it** - the worst kind of incident, because every dashboard says nothing happened. The fix is structural: storage keyed by content digest that rejects a rewrite, a digest check when serving loads the artifact, and a new registered version for every new set of bytes.
go deeper
Remember that a registered model version must point at fixed bytes. If a later job can write over the same location, the version number no longer tells you which model is running.
Explain content addressing: the run hashes the artifact, the store is keyed by that hash, and the registry holds the hash. A rewrite then becomes impossible rather than merely discouraged.
Reason about the incident shape - instances recycling into different bytes under one unchanged version, no deployment in the history, inputs steady - and name the controls: digest verification at load and reconciliation of loaded digests against the registry.
Balance the storage bill against evidence. Reference-driven retention, withdrawal instead of deletion, and an appended correction instead of an edited metric are the policies that keep past decisions defensible without keeping everything forever.
## What a model version is supposed to mean A registered model version is a promise: *these exact bytes, produced by that run, evaluated at that number, approved by that person*. The promise holds only if the bytes cannot change after the fact. The moment an artifact is addressed by a path that a later job can write to again, the version stops naming a model and starts naming a location. The practical form is **content addressing**: the run computes a digest of the artifact (SHA-256, for instance), the artifact store is keyed by that digest, and the registry entry holds the digest rather than a path. Two consequences follow for free - a rewrite is impossible because different bytes have a different key, and any reader can verify what it loaded. ## The four failures of an in-place overwrite 1. **The metrics become a lie.** The version's recorded evaluation error was measured on the previous bytes. Nothing marks it stale, so the number is still quoted in reviews. 2. **The approval becomes void.** Whoever promoted the version approved a specific model on specific evidence. After the overwrite, an approval record exists for bytes nobody reviewed. 3. **The rollback target may be gone.** If the overwritten artifact was the previously promoted version, the fallback for the next incident has been destroyed by a routine rerun. 4. **Serving becomes non-deterministic across instances.** Instances load the artifact when they start. After an overwrite, the fleet is split: instances started before the overwrite serve one model, those started after serve another, and an autoscaling event silently changes the mix. Failure four is the one that produces the memorable incident. Predicted claim severity drifts upward over a few days as instances recycle; the deployment history is empty; the model version in every log line is unchanged; drift dashboards on the inputs show nothing, because the inputs did not move. ## Immutable, superseded, withdrawn | situation | wrong move | right move | |---|---|---| | the model must change | overwrite the artifact | register a new version and promote it | | the model was harmful | delete the artifact | mark the version withdrawn, keep the bytes and the record | | a metric was computed wrongly | edit the recorded value | append a corrected evaluation, keeping the original visible | | storage is expensive | delete by age | delete by reference, sparing anything a version or approval points at | The common thread: evidence is appended to, never mutated and never erased. A past decision must still be explainable even when it was the wrong decision - especially then. ## Where the digest has to be checked Computing a digest on write and never checking it again gives you an identifier, not an integrity guarantee. The check that matters is **at load**: the serving process fetches the bytes for the promoted version, hashes them, and refuses to serve if the digest does not match the registry entry. A refusal to start is a good outcome here - a loud failure in one instance instead of a quiet, unattributable prediction shift across a fleet. The same check belongs in any step that copies artifacts between environments, which is where truncation and partial writes actually happen. ## What immutability costs, and the honest limit Immutability is not free: every candidate of every quarter keeps its bytes, and for large models that is a real bill. The answer is **reference-driven retention** rather than age-driven deletion: - any digest referenced by a registered version, a promotion record, the currently promoted version or the one before it is retained; - unreferenced candidate artifacts age out on a stated schedule; - deletion of a referenced digest is not a policy decision but a defect, and the store should refuse it. The honest limit is that immutability of the artifact alone does not make the served behaviour immutable. The same bytes fed different feature definitions produce different predictions, which is why a version entry has to pin its input contract as well as its digest. Immutable bytes are the floor, not the whole guarantee.
- Can a model version's artifact ever legitimately be replaced?No - it is superseded instead. If the bytes must change, register a new version and promote it; if the old one was harmful, mark that version withdrawn so it can never be promoted again, while keeping the bytes and its record so the original decision stays explainable. Mutating evidence and deleting evidence fail in the same way.
- How would you detect that an overwrite has already happened?Verify digests at load and on every copy between environments, and record the digest actually loaded in each serving instance's startup log. Reconciling loaded digests against the registry entry for the promoted version turns a silent fleet split into an alert. Without that reconciliation the only symptom is a prediction shift with no deployment behind it.
saying these in an interview costs you the question
- Stores the artifact under a path named latest and calls it a version
- Says overwriting is harmless because the metrics were already recorded
- Verifies the artifact digest on write but never on load
- Deletes artifacts on a size policy without checking registry references
- Assumes a version number alone identifies which bytes are served
- Fixes a bad model by replacing the bytes of the promoted version