skip to content

After the source rotates a ledger's datastore credential, which copies of the old value are still in use, and what makes the cutover survivable?

level: seniorimportance: should knowfreq 47%

answer

  1. issuing is not adopting
  2. reaches nothing already read
  3. two values accepted on purpose
  4. usage of the old value is the progress bar
  5. close the window, then actually revoke

basics

~20 s

Rotating at the source only changes what it hands out next; every copy already read by a running process still holds the old value. The cutover survives when the verifier accepts old and new together for a bounded window, so holders can converge in any order.

solid answer

~50 s

Issuing a new value is not the same as the fleet using it. Rotation reaches whatever asks the source *after* the change; it does not reach the copy an instance read at start-up, a value delivered into an instance that is still running, a value inside a shipped artifact, or copies held outside the fleet — a batch job, a reporting tool, an operator's machine, a partner. So a rotation without a plan fails in two ways at once: some holders break immediately, and others keep working on a value you believe is gone. The mechanism that makes it survivable is a **dual-accept window**: the verifier accepts the old and the new value at the same time for a bounded period, so each holder can pick the new one up whenever it next restarts or renews. You close the window only after the old value has gone unused for longer than the slowest holder's refresh interval — and then you actually retire it.

code

pseudocode · 12 lines
pseudocode
# at the verifier, during the rotation window
accept(presented):
    if matches(presented, current):
        return ok
    if matches(presented, previous) and now < previous.retire_at:
        record_use(previous, caller_id)    # each hit is a holder not yet moved
        return ok
    return reject

# closing the window
safe_to_retire(previous):
    return silence_since(record_use(previous)) > slowest_holder_interval

go deeper

for a junior

Understand that a credential can be changed where it is issued while running processes carry on with the value they read earlier; those are two separate events.

for a middle

List the places an old value still sits — process memory, a still-running instance, a shipped artifact, an occasional job — and say what replaces each one.

for a senior

Run the cutover: issue, accept both, let holders converge, watch usage of the old value to find the ones you missed, then close the window and revoke.

for a principal

Make rotation a routine operation rather than an incident: require dual-accept support when choosing what the fleet authenticates to, and treat an unrotatable credential as a design defect to be funded.

## Rotation at the source is not rotation in the fleet "Rotated" is one of the most overloaded words in this area. At the **source**, rotation means the credential store now issues a different value. At the **verifier** — the datastore, or the partner on the other end of an outbound call — it means which values are accepted. In the **fleet**, it means which value each running process is actually presenting. Those three move at different times, and every rotation incident lives in the gap between them. The rule to state plainly: rotating at the source changes what the source hands out next. It reaches no copy that a process has already read. ## Where the old value still is Work through the copies in a fleet that has already been delivered to: - **In the memory of every running instance**, read once at start-up. This is the big one, and it lasts as long as the instance does. - **In whatever was delivered into an instance that is still running.** If the platform replaces a delivered file, the bytes on disk may be new while the process is still using what it read at start — those are two different copies, and only the second one matters to the verifier. - **Inside any artifact that shipped with the value in it**, on every host and in every registry copy of that artifact. - **Outside the fleet entirely**: a nightly batch job that runs on its own schedule, a reporting tool, an analytics connector, an operator's machine, a runbook, a ticket attachment, and — for an outbound partner key — the partner's own systems. | Copy of the old value | Does rotating at the source change it? | What actually replaces it | | --- | --- | --- | | Held in a running process's memory | No | A restart, or an in-process renewal | | Delivered into a still-running instance | The bytes may change; the copy in use does not | Replacing the instance, or a re-read | | Inside a shipped artifact | No | A new artifact and a rollout | | Held by a job that runs weekly | No, until it next starts | Its next run, which may be days away | | Held outside your estate | No | A coordinated hand-over with the holder | A worked case. The source issues a new datastore credential at 10:00. The ledger's instances started at 08:00 and read the old value then. Nothing breaks at 10:00 — they are not re-reading anything. If the old value is revoked at the same moment, every instance breaks at its next *connection attempt*, which may be minutes later under pool churn or hours later when an unrelated deployment restarts them. That delayed, seemingly unrelated failure is the signature of a rotation with no window. ## The dual-accept window The mechanism that makes a cutover survivable is to let the verifier accept two values at once for a bounded time: 1. **Issue** the new value while the old one still works. 2. **Accept both** at the verifier; now the order in which holders converge stops mattering. 3. **Let holders converge**: fetchers pick it up at their next renewal, delivered holders at their next replacement, the weekly job at its next run. 4. **Watch the old value's usage.** Every acceptance of the old value is a holder you have not found yet, and that signal is the real progress bar for the rotation. 5. **Retire the old value** once it has gone unused for longer than the slowest holder's interval — and retire it for real, because an old value that is merely unused is still a valid credential. Step 4 is the step that gets skipped, and it is the one that turns rotation from a hopeful broadcast into a measurable operation. Step 5 is the one that makes it worth doing: a rotation that never retires the old value has doubled the number of live credentials rather than replaced one. ## When you cannot have a window Some verifiers accept exactly one value at a time — outbound partner keys are often like this. Then the cutover is genuinely atomic and you have two honest options: hold a second identity at the partner so that two values legitimately exist and you can move traffic between them, or schedule a real cutover with everything that implies. What you should not do is pretend the atomic cutover is routine and run it under load without a rehearsal. ## The judgement the question is testing Rotation is not measured by the moment the new value is issued. It is measured by the moment the last holder of the old value is gone and the old value has been revoked. Between those two moments the system is running on two credentials on purpose, and the engineering is in knowing who still holds which, not in the issuing.

  • Why does a rotation with no window often break a service hours after the cutover rather than at it?
    Because credentials are checked when a connection is established, not on every request. An instance that authenticated at start-up keeps working on an already-open connection, so the rejection first appears at the next reconnect — pool churn, a scaling event, or an unrelated restart. The gap between cause and symptom is what makes these incidents hard to attribute.
  • How do you find holders of the old value that nobody remembers?
    Instrument acceptance of the old value at the verifier and record who presented it. During the window every acceptance is a holder you have not migrated, with a caller identity attached. That turns the search from an archaeology exercise across repositories and runbooks into a shrinking list you can watch reach zero.
  • The old value has shown no usage for a week. Is the rotation finished?
    Not until it is revoked. An unused credential is still a valid one, and it is still sitting in whatever copies you never found. "No longer used" and "no longer works" are different states, and only the second one shrinks the blast radius of the copies still out there.

Changing the lock at the front door does not collect the copies of the old key already on people's keyrings. The dual-accept window is the period when both keys turn, and you only retire the old one once nobody has used it for a while.

saying these in an interview costs you the question

  • Thinks issuing a new value at the source rotates the fleet
  • Assumes a rotation that breaks nothing at cutover has succeeded
  • Forgets holders outside the fleet, like a weekly batch job
  • Leaves the old value valid because nothing seems to use it
  • Expects a delivered file's replacement to change what the process is presenting
  • Plans the cutover as a simultaneous restart of everything