skip to content

Schedules & Runbooks

Rotation the store performs on a schedule against a runbook a person follows at three in the morning: what the downstream must support, and the consumers nobody told. Asked because it fails quietly.

on this pageshow

questions

5

A store rotating a downstream credential on a schedule replaces a person working a runbook — which failures does that remove, and which does it add?

level: middleimportance: must knowfreq 62%

answer

  1. two lists, not one
  2. what a tired person gets wrong
  3. unattended means unwatched
  4. a clock knows nothing about context
  5. the rotator is itself privileged

basics

~20 s

Scheduling removes the failures of the human step — skipped runs, mistyped values, steps out of order — and adds silent failure, a clock blind to context, and a standing privileged holder able to change credentials.

solid answer

~50 s

Two different lists. A schedule removes everything that goes wrong because a person is doing it: the run nobody had time for, the value transcribed between two windows, the step taken out of order at three in the morning, and the need for anyone to see the value at all. It adds three things. The run is unattended, so its failure is silent until a consumer breaks — which means the schedule now needs its own alert and a check on the age of the value. The clock knows nothing about change freezes, peak load, or whether the downstream is reachable. And the rotating party becomes a standing holder of the right to change that credential, so it needs narrow rights and its own replacement path. A back-out, an owner and a stated done-signal stay written down either way.

go deeper

for a junior

Know that a credential in use can be replaced on a timer by the system that holds it, or by a person following written steps, and that both are production changes with their own ways to go wrong.

for a middle

Explain the two failure lists side by side: what the human step gets wrong, and what unattended execution hides. Name the alert that catches a scheduled run which has quietly stopped changing anything.

for a senior

Show that you have operated one. Talk about the done-signal, the rights the rotating party holds, the change-freeze decision, and why you would not automate a sequence the team has never performed by hand.

for a principal

Automation is a build-and-own decision per integration. Weigh the exposure it removes against the cost of the path, the monitoring it obliges you to run for ever, and the standing privileged holder it creates.

## Two shapes for replacing a value already in use A credential that something is currently authenticating with can be replaced in one of two shapes. In **scheduled, store-driven rotation** a timer fires, something holding rights at the downstream generates a new value and applies it there, and the new value is recorded where consumers read it. No person is in the path, and in the better designs no person ever sees the value. In a **runbook** a person works through a written sequence, usually because the downstream will not accept a change any other way — a supplier portal, an approval, a ticket queue — or because the sequence crosses systems nothing has been given rights over. Stores differ in how much of this they will do for you: some will perform the downstream change themselves against systems they know how to talk to, others only hold whatever value you hand them and leave the change to you. Treat that as a property of the store in front of you, not as a fact about stores. The interview question is never which shape is better. It is which failures each one removes and which it creates, because those are two different lists, and a team that automates without reading the second one has traded a noisy failure for a quiet one. ## The failures a schedule removes - **The run that never happened.** A change owned by a person competes with everything else that person owns, and loses. - **Transcription.** Any step where a human copies a value from one window into another is a step where it can be truncated, wrapped, or pasted into the wrong field. - **Order.** A sequence executed at three in the morning is executed by someone tired; a schedule performs the same order every time, correct or not. - **Visibility of the value.** A generated value can be applied downstream and recorded without ever being displayed, which removes a whole class of copies before they exist. - **Drift between the document and reality.** A runbook rots in silence; a schedule that has stopped working at least leaves a run history behind. ## The failures a schedule adds - **Silence.** This is the defining one. An unattended run has no observer, so a failure — or worse, a run that reports success while doing nothing — is discovered by whichever consumer breaks first, days or weeks later. - **A clock with no context.** The schedule does not know about a change freeze, a peak hour, a downstream in maintenance, or the release going out at that moment. - **Half a pair.** Automation can only touch what the downstream exposes. It is entirely ordinary for a schedule to change the side it can change and leave the other side to a person nobody assigned. - **A standing privileged holder.** Something now holds the right to change that credential all the time, rather than for the ten minutes a person held it. Its rights should be narrow, its own credential is long-lived, and it is worth attacking. - **Lost rehearsal.** After a year of automation nobody on the team has performed the change by hand, which matters on exactly the day the automation is the broken thing. | | Scheduled run | Runbook | |---|---|---| | Executed by | a timer and a privileged identity | a named person | | Typical failure | silent: a run that fails or quietly no-ops | skipped, mis-ordered, or half-completed | | How you find out | only if something watches the run and the value's age | the person tells you, or nobody does | | Prerequisite | a programmatic change path at the downstream | a person available inside the change window | | Cost per cycle | monitoring and ownership of the path | human time, in the worst hours | ## What neither shape decides for you Four things stay yours whichever shape you choose: 1. **A done-signal.** Something observable that says the change is complete — not `exit 0`, but evidence that the accepting side is on the new value and that a consumer has authenticated with it. 2. **An owner.** A schedule is not an owner. Someone is accountable for the alert when it stops, and that name belongs somewhere findable. 3. **A back-out.** What happens when the new value does not work, written as steps rather than as an intention. 4. **A record.** What changed, when, and by which path — the thing an external review asks about long afterwards. ## The direction people state backwards **Rotation puts a new value in place; it does not make the old one stop working.** Withdrawal is a separate action against the system that accepts the credential, and neither a schedule nor a runbook performs it unless it was written in. A rotation that ran flawlessly every night and never withdrew anything has bounded nothing at all — it has simply produced a series of values, one of which is in use.

  • The scheduled rotation has run nightly for a year. What single check tells you it is still doing something?
    The age of the value in use. Compare the newest recorded version against what the downstream last accepted, and alert when either stops advancing. A green run history proves the job executed, not that anything changed; a value whose age stops moving is the signal that survives a job which exits successfully on a no-op.
  • Your team automates a rotation nobody has ever performed by hand. What is the risk?
    The automation encodes a sequence no one has validated. The first time it half-fails, nobody has lived knowledge of the correct order, the verification step or the back-out, and the recovery is being invented under pressure. Automate the sequence you have already executed and could still execute; a path with no manual fallback has no fallback.
  • Does a change freeze apply to a scheduled rotation?
    It has to be decided explicitly, because the schedule cannot decide it. A rotation is a production change with an outage mode, so either the schedule is suspended for the freeze and the suspension is tracked to a resumed run, or the rotation is declared exempt and the reason recorded. Firing silently through a freeze is how a rotation becomes an unexplained incident.

saying these in an interview costs you the question

  • Automated rotation means the runbook can be deleted
  • No alert from the job means the rotation succeeded
  • Writing the new value also stops the old one working
  • Automate every credential, whatever the downstream supports
  • A runbook is safe because it is written down
  • The scheduler is safer than a person, so give it broad rights
open as a page

Your runbook put a new partner credential in the store and filed the supplier's change ticket; the nightly export broke six weeks later — what did that run leave undone?

level: seniorimportance: must knowfreq 58%

basics

~20 s

The run was closed when the store held the new value and the ticket was filed — neither of which changes what the export presents. Nothing moved the consumer, and nothing waited for the supplier's change, which landed six weeks later.

open as a page

Before a store can rotate a partner system's credential on a schedule, what must that downstream system itself support?

level: middleimportance: should knowfreq 50%

basics

~20 s

Four things from the downstream: a programmatic way to change the credential, an identity allowed to make that change, rules the generated value can satisfy, and a way to test the new value afterwards. Absence of any one forces a runbook.

open as a page

Half your estate's downstreams accept a programmatic credential change and half only through a supplier portal — how do you decide where automation goes?

level: principalimportance: should knowfreq 36%

basics

~20 s

Rank by what an older value costs, not by what automates easily. Build one shared path for the programmatic half; for the portal half, buy the exposure down with narrower rights, fewer holders and a rehearsed runbook.

open as a page

A scheduled rotation changed a datastore account's password but crashed before recording the new value — why is re-running it not automatically safe?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

The live value is now lost and the two sides disagree. A re-run may be unable to authenticate to make a second change, may consume the account's only spare credential slot, or may trip lockout and reuse rules. Reconcile first.

open as a page