skip to content

Your runbook put a new partner credential in the store and filed the supplier's change ticket; the nightly export broke six weeks later — what did that run leave undone?

level: seniorimportance: must knowfreq 58%

answer

  1. two halves, one clock
  2. written is not used
  3. who authenticated with the new value
  4. the supplier's queue owns the cutover
  5. done means the accepting side changed

basics

~20 s

The run was closed when the store held the new value and the ticket was filed — neither of which changes what the export presents. Nothing moved the consumer, and nothing waited for the supplier's change, which landed six weeks later.

solid answer

~50 s

Two halves of one change were fired independently and neither was verified. Recording a value in a store does not make any consumer use it: the export reads its credential when it starts and had not restarted since March, so it kept presenting the old value. That worked for six weeks only because the supplier had not processed the ticket. When they did, the old value stopped authenticating and the value that would have worked had been sitting unread in the store the whole time. The run's done-signal was "I performed my steps". A rotation is done when the accepting side is on the new value and a consumer has authenticated with it — and when the accepting side's change lands on someone else's clock, the runbook owns the wait, the check and the back-out, not just the request.

code

json · 14 lines
json
{
  "rotationRun": {
    "startedAt": "2026-03-04T01:00:00Z",
    "account": "partner-export",
    "storeWrite": { "status": "ok", "newVersionRecorded": true },
    "downstreamChange": {
      "method": "supplier-ticket",
      "status": "requested",
      "appliedAt": null
    },
    "consumersConfirmedOnNewValue": [],
    "closedAs": "success"
  }
}

go deeper

for a junior

Know that putting a new credential into a store does not make anything start using it, and that a rotation is not finished until the system accepting the credential has the new value too.

for a middle

Separate the three facts: what the store would return, what the consumer is presenting, and what the partner will accept. Explain why the run's success record spoke to none of the last two.

for a senior

Own the interval. Say what you would alert on between the store write and the supplier's change, what evidence closes the run, and what the back-out is when the change lands at 02:00.

for a principal

Decide what 'rotated' means for the whole estate and make the completion criterion the reported one. A programme measured by runs performed will keep producing this outage, quietly, everywhere.

## What actually happened 1. A person generated a new credential for the partner integration and recorded it in the store. That step succeeded. 2. The same person filed the supplier's change request, because the partner changes credentials only through a portal and a ticket queue. The run was closed green that morning. 3. The nightly export read its credential when its process last started, in March. It did not restart, so it kept presenting the **old** value — which still worked, because nothing had changed at the partner. 4. For six weeks the store held a value that had never been read by anything and would not have authenticated if it had been. 5. The supplier applied the change. At 02:00 the next night the export failed to authenticate, and the first news of a six-week-old rotation was an outage. ## Why "the store holds the new value" is not a done-signal A store write changes what a **read would return**. It does not change what a running consumer is presenting, and it does not change what the accepting system will accept. Those are three separate facts, and only the third one decides whether the export works. - The store's record answers *what would I be handed now*. - The consumer's behaviour answers *what am I presenting*. - The partner's configuration answers *what will be accepted*. A rotation is finished when the second and third agree on the new value, and the old value has been withdrawn. Nothing in this run measured either. ## The half that lands on someone else's clock The hardest property of this shape of change is that **you do not control when the second half happens**. A ticket in a supplier's queue lands in a window measured in weeks, and the interval between your write and their change belongs to nobody by default. Two directions matter here and both were left open: - **Before** the supplier's change, the new value is inert. Anything that started using it early would have failed immediately. - **After** the supplier's change, the old value is withdrawn whether or not any consumer moved. Their change is simultaneously the replacement and the withdrawal, because the partner holds one credential for this integration. Ordering the store write first was reasonable — it means the value is on record and recoverable before the cutover. It becomes a defect only because the run was closed at that point instead of staying open across the gap. | The run's record said | The run's record never said | |---|---| | a new version was written | whether anything had read it | | a change request was filed | whether the supplier had applied it | | every documented step was performed | which value the export was presenting | | the operator who ran it | when the change would become live | ## What a done-signal looks like here - **Evidence from the accepting side**: a confirmed change date, or an observed authentication with the new value. - **Evidence from a consumer**: at least one consumer seen authenticating successfully after the change, which is what "moved across" means in practice. - **An alert on the gap**: the expected change date passes without confirmation, or the newest recorded value has never been read. - **A back-out**: what happens if the new value does not work when the supplier applies it, at 02:00, with the export already failing. ## What this scenario is not about Two neighbouring subjects are deliberately left alone. *How* a changed value reaches a running process — caching, refresh, restart — is a delivery question with its own answer. *Who* all the holders are is a census exercise of its own. The defect here is narrower and more common than either: a rotation whose definition of done was the set of steps the operator performed, in a change whose decisive event happens on another organisation's schedule. The general rule is worth stating plainly, because it survives every store design: **a rotation that nothing has used is not a rotation, it is a stored value.** Until something authenticates with it, the credential in production is still the old one, and the risk the rotation was supposed to reduce has not moved at all.

  • What would you have watched during the six-week wait?
    The gap itself. Record the expected change date and alert when it passes with no confirmation, and alert when the newest recorded value has never been read by anything. Both signals existed for six weeks and neither was wired up, which is why an authentication failure at 02:00 was the first news of the change.
  • The supplier's change can land any time in a six-week window. How do you order it?
    Stop treating the request as the last step. Either obtain a scheduled date and hold the run open until it passes, or ask the partner for a second credential so the new value is live before the old one goes. Where they hold only one, the honest answer is a booked change window and someone watching it — not a run closed the morning the ticket was filed.
  • Whose failure is it, when the run reported success?
    The definition of done. A record of which steps were performed answers a different question from a record of which value is in use, and only the second could have been checked at any point during those six weeks. Fix the completion criterion and the same runbook stops producing this outage.

saying these in an interview costs you the question

  • The rotation succeeded because the store holds the new value
  • Filing the supplier's change request completes the rotation
  • The export would have picked the new value up by itself
  • A six-week gap is harmless because the old value still worked
  • Nothing could be done while waiting for the supplier
  • The outage proves the new value was generated wrongly