skip to content

You are moving to a new case-management product and the old one holds years of recorded execution history. How do you decide whether to carry that history across, and what do you do if you do not?

level: principalimportance: should knowfreq 40%

answer

  1. ask who actually reads old runs
  2. provenance is what does not travel
  3. keep history outside the new product
  4. a stated seam beats a hidden gap
  5. set the retirement date now

basics

~20 s

Decide by who actually queries it. Most teams read only recent history, so migrate definitions and keep history outside the new product: the old instance read-only for a fixed window, then a neutral file export that stays searchable.

solid answer

~50 s

Start from the demand side, not the data side: find out who genuinely reads old executions and for what. In most teams the honest answer is a handful of people, a few times a year, asking whether something ever passed or reconstructing what happened around a defect. That demand rarely justifies a faithful replay, which is the expensive half of a migration — timestamps, the original runner and the surrounding cycle usually cannot be recreated in the target, so what you get is history that looks complete and is not. Prefer keeping history *outside* the new product: leave the old instance read-only for a stated window, then take a neutral export of runs and their evidence, keyed by the old case identifier, into storage the team can still search. If an external obligation applies, that constraint is a floor set elsewhere, not a preference. Above all, avoid a partial load — a report over half-migrated history is confidently wrong.

go deeper

for a junior

Know that a migration normally carries case definitions and not past runs, and that anything you want to keep from the old system has to be decided on before that system is switched off.

for a middle

Explain why replaying history is the hard half: timestamps and the people who ran things cannot be genuinely recreated, the surrounding grouping has no counterpart, and attached evidence is usually the first thing skipped.

for a senior

Show the practical alternatives and their operational cost — a read-only source for a stated window, a bounded carry with a visible seam, or a neutral keyed export — and insist the old identifier is preserved so any archive stays findable.

for a principal

Own the trade and write it down: what demand exists, which floor is set externally, where the seam falls, who owns the archive, and the date the old instance dies. An unstated decision here becomes a permanent second system.

## Start with the demand, not the data The instinctive framing is *how do we move the history?* The useful framing is *who reads it, how often, and to answer what?* Almost every migration plan is improved by asking a few people that question before anyone estimates the work. The honest answers usually cluster into three: - **Has this ever passed?** Someone arguing about a shaky area wants to know whether a case has a track record. They want recent history, not years of it. - **What happened around this defect?** Archaeology on a specific incident, needing a bounded window and usually the evidence attached to it. - **Show me the record.** Someone outside the team who needs to see that something was verified. This is the demand that can genuinely require completeness — and where an obligation exists, it is a floor set outside the migration, not a preference to trade off. If nobody names a fourth, you have just discovered that most of the history is being kept out of unease rather than use. ## Why history is the expensive half Definitions are text and structure; runs are relationships and provenance, and provenance is what does not travel: - **Time cannot be honestly recreated.** A loaded run happens now, not when it happened. Where a product accepts a supplied date at all, it is a value you asserted rather than a fact it observed. - **People cannot be recreated.** The person who ran it may not exist as an account in the target, and mapping leavers onto a placeholder rewrites the record while appearing to preserve it. - **The surrounding context has to be invented.** A run belongs to something — a grouping the old product created and the new one has no counterpart for — so the load either fabricates containers or drops the association. - **Evidence multiplies the cost.** Files attached to old runs are the bulk of the volume and the part most likely to be skipped, so the loaded history is often the outcomes without the proof. The result is that a full replay is expensive, slow, and *still* lower fidelity than the source it came from. ## Four options, and what each really costs | Option | What you get | What it costs | |---|---|---| | Full replay into the target | History visible where people now work | The highest effort, and provenance you asserted rather than preserved | | Bounded carry (a recent window, or the release still supported) | The queries people actually make are answerable in one place | A visible seam, and the discipline to state where history begins | | Old instance kept read-only | Perfect fidelity, no load work at all | Two systems to reach, and a running cost until the stated end date | | Neutral export to files | Durable, cheap to keep, vendor-independent | Not queryable the way a product is; needs a key and an index to be usable | The fourth deserves more credit than it usually gets. A per-case export of runs and their evidence, keyed by the **old** case identifier and stored where the team already searches, outlives every product involved and costs almost nothing to hold. It is a poor daily tool and an excellent archive. ## How to decide 1. **Count the real queries.** How many times last year did anyone open an execution older than the last couple of releases? Ask, do not guess. 2. **Separate obligation from habit.** Where an external party must be able to see the record, that requirement sets a floor and is decided outside this migration. Everything above the floor is yours to trade. 3. **Price the seam.** A bounded carry means the new product's history starts on a date. That is fine if the date is stated and everyone knows where to look for earlier records — and corrosive if it is discovered. 4. **Check the key exists.** Any archive is worthless without the old case identifier recorded on the migrated cases. If you have not parked the source key during the load, nothing you keep can be found again. 5. **Set the retirement date at the same time.** An old instance kept alive with no end date is the default outcome and the most expensive one. Decide the date while everyone is still paying attention. ## The trap: half-migrated history The worst outcome is a partial load nobody labelled. It looks complete. Reports over it are computed over a subset that reflects what the importer managed rather than what happened, and someone will make a decision on those numbers. A gap you can see is a caveat; a gap you cannot is a lie. So pick an end of the spectrum deliberately. Either the new product's history starts clean at cutover, with the old record preserved somewhere stated, or it is genuinely complete. What must not happen is landing between the two by accident and letting the reporting surface treat the result as whole. ## What good looks like A one-page decision, written down: history in the new product begins at cutover; the old instance stays readable until a named date; before that date a neutral export of runs and evidence, keyed by the old identifier, goes into named storage with a named owner; and anyone needing an older record knows exactly where to go. Written down, that is a decision. Left implicit, it is the thing someone discovers a year later when the old instance has already been switched off.

  • If you keep the old instance read-only, what has to be true for that to still work a year later?
    A named owner, a stated end date, and access that does not silently lapse — accounts, the gateway in front of it, and whatever certificate or login path it depends on. Read-only systems rot quietly because nobody uses them until the day they are needed. Put the end date and the owner in the same document as the migration decision.
  • Someone insists the whole history must be visible inside the new product. How do you handle that?
    Ask what question they need answered and how often. If it is an external obligation, that is a floor and I plan to it. If it is comfort, I show what a replay actually produces — asserted timestamps, remapped people, missing evidence — and offer the archive instead, which is more faithful than the loaded version they were picturing.

saying these in an interview costs you the question

  • Assumes every past run must land in the new product
  • Loads part of the history without labelling the gap
  • Treats a supplied timestamp as preserved provenance
  • Keeps the old instance alive with no end date
  • Archives runs without recording the old case identifier