skip to content

An operations tool offers a dry-run mode that prints the changes it would make without making them. What does running that mode actually prove before you let the tool loose on production, and what can it still not tell you?

level: juniorimportance: nice to knowfreq 40%

answer

  1. it previews scope, not success
  2. count the targets before confirming
  3. the write path is never exercised
  4. the plan is stale the moment it prints
  5. dry-run code can differ from real code

basics

~20 s

A dry run proves the tool's inputs, targeting and scope: which objects it selected and how many it would change. It cannot prove the change succeeds, that permissions allow it, that the world will still look the same at run time, or that a partial failure is safe.

solid answer

~50 s

The most valuable thing a dry run gives you is a blast-radius preview. You wanted to touch four things; the output says it will touch four thousand. That single check catches the classic operations disaster — a selector or filter that matched far more than intended — before it is irreversible. What it cannot tell you is anything about the write path. Permissions are usually only exercised on the real call, so a dry run can pass while the actual change is rejected halfway through. It says nothing about ordering or about what state you are left in if the tool dies on item 300 of 1,000. And the state it inspected is already stale by the time you confirm — something can change in between. There is also a subtle trap: if the dry-run branch is separate code, it can be wrong about what the real path does.

go deeper

for a junior

Be able to say that a dry run shows what would change and how many things it would touch, and that a clean dry run is not a guarantee the real run will succeed.

for a middle

Explain the specific gaps — the write path and its permissions are never exercised, downstream validation happens only on apply, and the plan goes stale between preview and confirmation.

for a senior

Talk about the preview being trustworthy only if it shares a code path with the apply, about narrowing the preview-to-apply window, and about proving the write path on a small real subset before a bulk change.

for a principal

Make the preview a required review artefact for any automation with destructive authority, and require that the plan a reviewer approves is the plan that executes rather than one recomputed later.

## What a dry run is for A dry-run (sometimes "preview", "what-if" or "no-op") mode executes the tool's decision-making but suppresses its side effects. It is the cheapest safety mechanism in operations, and it earns its place mainly by answering one question: **how much is this about to touch?** The classic operations disaster is a scope error, not a logic error. A selector that was meant to match one deployment matches every deployment in the namespace. A cleanup that was meant to delete objects older than 90 days parses the date wrong and matches everything. A configuration push aimed at the canary group is aimed at the whole fleet. In every one of those cases the tool did exactly what it was told; the operator was wrong about what they told it. Reading a count and a sample of the targets before confirming catches all of them. ## What it genuinely verifies - **Input parsing and resolution.** The arguments, config file and environment resolved to the values you meant. - **Target selection.** Which objects matched, and how many. - **The planned action per target.** Create, update, replace, delete — and "replace" or "delete" appearing where you expected "update" is a stop signal. - **Read-path connectivity.** The tool could reach and authenticate against the system enough to enumerate state. ## What it structurally cannot verify - **That the write succeeds.** Read permission and write permission are different grants. A dry run that enumerates happily can be followed by a real run that is denied on the first mutation. - **Downstream effects.** Validation the target system performs only at write time, quota limits, admission rules, triggers, and anything the change causes in another system are all invisible. - **Partial-failure behaviour.** If the tool stops at item 300 of 1,000, what state is the system in? A dry run tells you nothing about this, and it is often the most dangerous property of a bulk operation. - **Freshness.** The plan describes the world as it was at read time. Between the preview and your confirmation, a deploy could land or an autoscaler could act. The narrower that window, the less it matters — and reviewing a plan for twenty minutes before approving it widens it considerably. - **Its own honesty.** If dry-run support is implemented as a separate branch (`if dry_run: print(...) else: apply(...)`), the two paths can diverge, and the printed plan is then a description of code that is not the code that will run. A dry run built by executing the same path with the effectful call stubbed at the boundary is far more trustworthy. ## Why it matters more the moment you automate For a human running a one-off command, a dry run is a courtesy. For anything scheduled or unattended, it becomes a development and review tool: the plan is what you inspect in staging, and what a reviewer reads before the automation is granted authority. Related to it — and often confused with it — is **idempotency**: a dry run tells you what *would* happen once, while idempotency is the property that makes running the real thing twice safe. Automation that will be retried needs both, because the retry is the normal case, not the exception. ## Practical discipline - Always read the *count* first, not the first few lines of output. - Treat any unexpected `delete` or `replace` in the plan as a hard stop, not a curiosity. - Keep the gap between preview and apply short. - On a large bulk change, run it against a small subset for real before trusting the plan for the whole set — a genuine write on ten objects proves more about the write path than a dry run on ten thousand. ## How to answer in an interview Name the blast-radius preview as the primary value, then be specific about at least two limits — the unexercised write path and the staleness of the plan. A candidate who says "dry run means it is safe" is telling the interviewer they have never had a dry run pass and the real run fail.

  • Why can a dry-run mode implemented as a separate code branch be misleading?
    Because the plan it prints describes the branch that ran, not the branch that will run. If the two drift — a condition added to the apply path but not the preview path — the preview is a confident description of the wrong behaviour. The safer implementation runs the same logic and stubs only the effectful call at the boundary, so there is one decision path.
  • You have a dry run that looks correct for a bulk change across 5,000 objects. What would you still do before running it fully?
    Run it for real against a small subset — ten or twenty objects — and verify the result. That exercises the write path, permissions, downstream validation and partial-failure behaviour, none of which the dry run touched. If the subset succeeds and looks right, proceed in batches rather than in one pass, so a failure part-way leaves a bounded and identifiable set of objects to reconcile.

saying these in an interview costs you the question

  • A clean dry run means the real run is safe
  • Dry run verifies that permissions are sufficient
  • It shows the plan, so the state cannot change before apply
  • Dry run and idempotency are the same guarantee
  • Skim the first lines instead of reading the target count

context