Re-encrypting a decade of records under a new key will take six weeks — how do you order minting, re-protection, retirement and confirmation, and what exposure do you accept meanwhile?
answer
- order the queue, not the calendar
- new writes first, cheapest step
- encrypting and decrypting are separate permissions
- two independent confirmations before retirement
- destruction is the irreversible one
basics
~20 sMint and redirect new writes first, keep the old key able to decrypt while data migrates, order the queue by what disclosure costs most, confirm from both the data side and the usage side, and only then retire — destruction last, because it is the irreversible step.
solid answer
~50 sThe order is forced by one constraint: the old key must stay able to **decrypt** until the last record has moved, or every un-migrated read fails. So mint the replacement and point new writes at it immediately — that is instant and stops the exposed set growing — then stop the old key **encrypting** while leaving decryption alone. The six weeks is a queue, and the only real lever is its order, so migrate by what disclosure would cost rather than by dataset size or storage order. Before retirement, confirm twice and independently: no stored ciphertext still names the old key, and its usage records show no operations for a defined quiet period. Destroy last, and only after the backup path no longer needs it, because destruction turns anything you missed into permanent data loss.
code
pseudocode · 20 linesnewKey = mint()
setEncryptingKey(dataset, newKey) // new writes stop joining the exposed set
allowDecryptOnly(suspectKey) // reads of un-migrated data must keep working
for each dataset in sortByDisclosureCost(descending):
reprotect(dataset, from: suspectKey, to: newKey)
// two independent confirmations, not one
dataSideClear = countCiphertext(keyId == suspectKey) == 0
usageSideClear = daysSince(lastOperation(suspectKey)) >= quietPeriod
backupPathClear = noRestorePointRequires(suspectKey)
if dataSideClear and usageSideClear:
withdrawDecrypt(suspectKey)
if backupPathClear:
destroy(suspectKey) // no way back from here
else:
holdDisabled(suspectKey) // restores still need it
else:
keepDecryptEnabled(suspectKey) // something unmigrated is still out therego deeper
Recall that the old key has to keep decrypting while data is still encrypted under it. Turning it off early breaks reads rather than protecting anything.
Explain that encrypting and decrypting are separate permissions, and that a migration uses that split: stop new writes on the old key immediately, keep reads working until the last record has moved.
Demonstrate the confirmations. No ciphertext still naming the old key, and no operations in the usage records for a period long enough to cover the rarest consumer — then retirement, with destruction handled separately.
Own the trade the queue makes. Six weeks of ordering decides which data sits exposed longest, and destruction against backup coverage is a loss you accept deliberately and record, not a step a runbook takes on its own.
## Why the order is forced Four actions, and the dependencies between them leave almost no freedom: 1. **Mint** the replacement and make it the encrypting key. Seconds of work, and from that moment the exposed set stops growing. Nothing justifies delaying it. 2. **Stop the old key encrypting**, while leaving it able to **decrypt**. These are two different permissions, and separating them is the whole trick of an orderly migration. 3. **Re-protect** the stored data — read with the old key, write back under the new one — dataset by dataset. 4. **Confirm, then retire**, and only afterwards consider destruction. The forced part is step 2. Withdrawing decryption early is the classic self-inflicted outage: every read of data that has not yet migrated fails, which on a decade of records means nearly all of it. The exposed key has to keep working, for you, for as long as the migration takes. ## What the six weeks actually costs Be precise about who the delay affects, because the instinct is to treat it as the exposure: - It does **not** affect the holder of the copy. They never needed your systems, your permissions or your uptime, and they were not waiting on your schedule. - It **does** affect the copies still inside your estate — backups, replicas, exports — which stay openable with the exposed key for as long as they exist under it. - It **does** extend the period in which somebody who gets into your systems could use the old key on data that has not moved yet. So the accepted exposure is a statement about your own copies and your own perimeter, not about the disclosure, which already completed. ## The queue's order is the decision Six weeks of work means six weeks of ordering, and the ordering decides who waits longest. Rank by **what disclosure of that dataset would cost**, not by: - **Dataset size** — finishing many small sets first makes a progress chart look good and protects the least valuable data first. - **Storage order** — a sequential pass is the fastest total runtime and the worst risk ordering. - **Recency** — new writes already use the replacement key, so recent data is not in the queue at all. A useful second axis is retention: a dataset whose backups keep the old ciphertext for years is worth moving before one whose copies age out in a fortnight. ## Confirming before you retire One check is not enough, because each one is blind to a different failure. Use two, independently: - **The data side**: no stored ciphertext anywhere still names the old `keyId`. This catches datasets the migration job never enumerated. - **The usage side**: the manager's usage records show no operations with the old key for a defined quiet period. This catches the consumer nobody documented, which the data-side check cannot see because it holds its own copy. Taking the migration job's own completion count as proof is the standard mistake — the job cannot report data it never knew about. | Step | Reversible? | What going too early costs | |---|---|---| | Mint replacement | Yes | Nothing; do it first | | Stop old key encrypting | Yes | Nothing meaningful | | Withdraw decryption | Usually yes | Every un-migrated read fails | | Destroy key material | **No** | Anything missed is unreadable forever | ## Destruction is the one you cannot take back Destruction is genuinely final, and two things routinely depend on the key after everyone believes it is finished: - **Backups taken during the window**, which open only with the old key. Destroy it and those restores are gone; the honest sequence is to re-take the backup path first, or to accept and document the loss of restore coverage for that period. - **Archived exports** somebody produced once and nobody catalogued. Some managers offer a disabled-but-recoverable state between retirement and destruction; some offer only deletion. Where the reversible state exists, use it as the soak period and destroy afterwards. Where it does not, the soak period has to be built out of the two confirmations above plus time. ## What to say when asked Give the order and the reason for each position: mint first because it is free, decryption last to leave because reads depend on it, the queue ordered by disclosure cost because the order is what decides who is exposed longest, two independent confirmations because each one is blind to the other's failure, and destruction after the backup path because it is the only step with no way back.
- What do you tell stakeholders about the six weeks, given the key is already in someone's backup?That the schedule does not change what the holder can read — they never depended on your systems. The window matters for copies still inside the estate, which stay openable with the exposed key until they are rewritten or expire. Report the exposure as the data the key covered, and the schedule as the plan for your own copies, so the two are not confused.
- Is there ever a case for destroying the exposed key before the migration finishes?Only where the reachable copies inside your estate are a bigger risk than losing the data — for example when the ciphertext is reproducible from a source of record, so nothing is truly lost. Otherwise destruction converts every missed record into permanent data loss, and it is the one step with no way back.
- How long should the quiet period on the usage records be?Long enough to cover the slowest thing that reads the data. If a quarterly report is the rarest consumer, a thirty-day quiet period proves nothing about it, and the period has to reach past that cycle. State the assumption explicitly: the period is a claim about how infrequently something might read, not a round number.
saying these in an interview costs you the question
- Withdraws decryption before the data has moved
- Destroys the exposed key while backups still need it
- Presents the re-encryption schedule as the remedy for the disclosure
- Confirms from the migration job's own completion count alone
- Orders the queue by dataset size rather than disclosure cost
- Treats retirement and destruction as the same action