skip to content

Toil, Runbooks & Automation

Identifying toil — manual, repetitive, automatable ops work — and eliminating it with runbooks and progressive automation. A favorite SRE interview theme because it distinguishes operators who scale themselves from those who just work harder.

on this pageshow

questions

16

In SRE, what makes operational work count as "toil", and name two kinds of manual work that are not toil?

level: juniorimportance: must knowfreq 72%

answer

  1. not every manual task qualifies
  2. six tests, roughly all must hold
  3. nothing enduring left behind
  4. grows as the service grows
  5. overhead and engineering are different buckets

basics

~20 s

Toil is manual, repetitive, automatable, tactical work with no enduring value that scales with the service. Manual work failing those tests — a one-off migration, or novel debugging of a new failure — is engineering or overhead, not toil.

solid answer

~50 s

Toil is a specific category, not a synonym for unglamorous work. The criteria are: it is **manual**, **repetitive**, **automatable**, **tactical** (reactive, interrupt-driven), it produces **no enduring value** — the service is in the same state afterwards as before — and it **scales linearly with service growth**. A task has to hit essentially all of those to count. So a one-off hardware migration is manual but not repetitive and leaves lasting value: that is project work. Debugging a novel failure mode is manual and reactive but not automatable, because you don't yet know what to encode: that is engineering. Team meetings, interviews and expense reports are overhead, not toil. The reason SRE bothers with a precise definition is budgeting: toil is the number you point at to justify spending engineering time on automation, and Google caps it at roughly 50% of an SRE's time. A definition loose enough to cover everything you dislike can't carry that argument.

go deeper

for a junior

Be able to state the definition in one sentence and list the criteria, and say plainly that toil is necessary work, not bad work — it is the share of time it consumes that is the problem.

for a middle

Explain why each criterion is there, especially automatable and no-enduring-value, and be ready to classify a borderline item such as a quarterly certificate renewal or a customer data-export request.

for a senior

Show that you use the definition as an instrument: point at a real ticket queue, say which fraction of it is toil, and defend the calls to a manager who wants the number smaller.

for a principal

Own the boundary across teams so the number means the same thing everywhere. A definition that drifts per team cannot be aggregated, and a loose one gets used to justify whatever automation someone already wanted to build.

## Why the word needs a definition at all Every operations team has work it resents, and every team would like that work automated. "Toil" is SRE's attempt to make that argument fundable rather than emotional. If toil is precisely defined and measured, a team can say "41% of our engineering time went to toil last quarter, here are the three sources, here is what we intend to spend to remove them" — and a manager can act on it. If toil just means "work we don't enjoy", the number is unfalsifiable and nobody will fund anything on the strength of it. ## The criteria Google's Site Reliability Engineering book gives six tests. Work is toil when it is: - **Manual** — a human has to touch it. Time spent hand-running a script counts; the script running itself does not. - **Repetitive** — you have done it before and you will do it again. First time is investigation; the twentieth time is toil. - **Automatable** — a machine could plausibly do it. If the task requires judgement no one has yet been able to encode, it is not toil, it is engineering. This is the criterion that keeps incident diagnosis out of the bucket. - **Tactical** — it is reactive and interrupt-driven rather than strategy-driven. A page you answer at 3am is tactical; the capacity plan you write on Tuesday is not. - **Devoid of enduring value** — when you are done, the service is in the same state it was in before. You restarted the stuck worker; nothing about tomorrow got better. - **Growing at least linearly with service size** — the cost of the task rises as the estate grows. This is what makes toil an existential problem rather than an annoyance: work that scales with the service eventually consumes the whole team. In practice you rarely need all six to make a call, but the two that do the most classifying work are *automatable* and *no enduring value*. If a task leaves something permanently better, or if nobody knows how a machine would do it, it is not toil. ## What is not toil Three categories get misfiled constantly. **Overhead.** Meetings, performance reviews, hiring loops, expense reports, planning. These are manual, repetitive and arguably valueless, but they are not operational work on your service, they do not scale with it, and your team cannot automate them away. They belong in a separate line of the time budget. Folding them into the toil number inflates it and makes the whole measurement easy to dismiss. **Project and engineering work.** A migration, a schema change, a new deployment pipeline. Manual, sometimes deeply tedious, but each one leaves the system permanently different. Doing it once is investment, not toil. If you find yourself doing the "one-off" migration for the eleventh time, it has become toil and you should say so. **Genuine incident response.** The novel outage where you don't know what is broken is the opposite of toil: it is not repetitive, not automatable, and the understanding it produces is enduring value. What *is* toil is the incident you have seen fifteen times, where you follow the same steps to the same resolution. ## Toil is not bad work, and it is not zero Two corollaries candidates often miss. First, toil is legitimate, necessary work. Someone genuinely does have to restart the worker. Calling it toil is not calling it beneath anyone, and an SRE who refuses to do operational work has misread the idea entirely. The claim is only that its *share* of the team's time has to be bounded, because a team at 100% toil has no capacity left to make next quarter cheaper than this one. Second, the target is not zero. Google's published guidance is a ceiling of roughly 50% of an SRE's time, precisely because some toil is the price of running a real service and because keeping engineers in touch with production has value. A team reporting 0% toil is either not operating anything or not measuring. ## Applying it The useful exercise in an interview is classification. Given a queue — quarterly certificate renewals, a customer's data-export request, a flaky test rerun, a capacity review, a manual production release, a novel latency investigation — say which are toil and why, and the honest answer will not be "all of them". Certificate renewal and the manual release are textbook toil: repetitive, automatable, leaving nothing behind, and each new service adds another one. The capacity review and the latency investigation are engineering. The export request is toil only if it keeps coming; if it is the first, it is a signal you may need a self-service feature. The classification is the whole point: it converts a vague complaint into a prioritised list of things worth building.

  • Is toil the same thing as operational work?
    No. Operational work includes incident response to novel failures, capacity reviews and design consultation — none of which are automatable or repetitive. Toil is the automatable, no-enduring-value subset. Teams that equate the two end up concluding that all ops work is bad, which is both wrong and corrosive to the relationship with the service owners.
  • If a task is manual and repetitive but takes five minutes a month, is it toil?
    Yes, by definition — but that says nothing about whether to automate it. The definition classifies; the measurement prioritises. Five minutes a month is a rounding error and probably stays manual forever. Being toil and being worth removing are separate judgements, and conflating them is how teams end up automating trivia while the real cost sits untouched.
  • Where do meetings, interviews and expense reports fit?
    Overhead, tracked as its own line. They fail the criteria on two counts: your team cannot automate them, and they do not scale with the service. Folding them into the toil number inflates it, and the first manager who spots that will discount the whole measurement — which costs you the argument you were trying to fund.

saying these in an interview costs you the question

  • Calls every manual or unglamorous task toil
  • Treats toil as work beneath senior engineers
  • Says the goal is zero toil
  • Counts meetings and admin as toil
  • Classifies novel incident debugging as toil
  • Thinks toil means the work is optional or unimportant

context

open as a page

A recurring production fix currently runs as a script an on-call engineer executes by hand after being paged. What do you gain, and what do you take on, by promoting it to closed-loop automation that runs with no human in the loop?

level: middleimportance: must knowfreq 60%

basics

~20 s

Closing the loop buys speed and consistency: the fix runs in seconds, identically, without waking anyone. In exchange you take on a new production actor that changes systems unsupervised, so it needs guardrails, its own telemetry, and a named owner.

open as a page

You are writing a runbook for a failure that pages the on-call engineer at 3am. What sections should it contain, and what does each section have to do for someone who did not build the service?

level: middleimportance: must knowfreq 68%

basics

~20 s

A usable runbook states its trigger and the user impact, then gives diagnostics with expected output, remediation with preconditions and blast radius, a verification signal that proves the fix worked, an escalation path with a time-box, and metadata naming the owner and last validation date.

open as a page

A watchdog restarts your API process whenever its health check fails. It has been quietly restarting the process about three times a night for two months; last night the restarts could not keep up and the service was down for 40 minutes. What went wrong in the way that automation was built?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The watchdog suppressed the symptom without reporting it, so a slowly worsening fault stayed invisible until it outran the remediation. Auto-remediation must count every action, escalate when the action rate climbs, and refuse to keep repairing indefinitely.

open as a page

Your team's runbooks have gone stale — commands reference a decommissioned host and several document alerts that no longer exist. How do you make runbook freshness a property of the process rather than a one-off cleanup project?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Tie maintenance to use: whoever follows a runbook during a page repairs it in the same shift. Bind each runbook to the alert that links it so orphans are visible, review them at on-call handoff, stamp owner and last-validated date, and delete more than you write.

open as a page

A team keeps a detailed architecture wiki and argues it does not need runbooks. What does a runbook give an on-caller at 3am that architecture documentation does not?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A runbook is a procedure for one known failure: the symptom it matches, the exact diagnostic and remediation steps, how to verify the fix, and when to escalate. Architecture documentation explains how the system works and leaves the responder to derive the response.

open as a page

Two recurring manual tasks each cost your team about an hour a week today. One's frequency grows in proportion to the number of services you run; the other is fixed no matter how large the estate gets. Why does SRE's definition of toil treat those differently, and which do you attack first?

level: middleimportance: should knowfreq 48%

basics

~20 s

Toil that scales with the service is the dangerous kind: its cost rises with the estate until it consumes the team. Attack the scaling task first — the fixed hour stays one hour forever, however irritating it is.

open as a page

A responder followed a runbook's "restart the service" step for a symptom that resembled the documented one but had a different cause, and made the outage worse. How prescriptive should a runbook be, and what has to guard a copy-paste command?

level: middleimportance: should knowfreq 40%

basics

~20 s

Be maximally prescriptive about mechanics and explicit about the conditions under which each step is valid. Every copy-paste command needs a stated precondition, its blast radius, whether it is reversible, and an escape hatch that says stop and escalate when the symptoms do not match.

open as a page

You are building automation that terminates and replaces any instance whose health check is failing. A monitoring bug briefly reports every instance in the fleet as unhealthy. What safeguards stop that from destroying the fleet?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Bound the automation before it acts: refuse to act when an implausible share of the fleet looks unhealthy, cap actions per time window, cap how much capacity may be missing at once, and trip a breaker that pages a human instead of continuing.

open as a page

Your manager asks what fraction of your team's time goes to toil. How do you actually measure that, and how do you keep the number honest enough to make decisions on?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Measure toil as a share of team engineering time per quarter, from ticketed interrupts, on-call logs and periodic sampling. Keep it honest by ticketing every interrupt, counting per team, and trusting the trend over the exact figure.

open as a page

A runbook's remediation has just been fully automated, so the fix now runs without a human. What should happen to the runbook itself?

level: seniorimportance: should knowfreq 34%

basics

~20 s

It changes role rather than disappearing: it now documents what the automation does, how to tell whether it ran, how to disable it, and the manual fallback for when it is off or broken. That fallback is never exercised, so it must be deliberately drilled or plainly marked unvalidated.

open as a page

Your team has a backlog of recurring manual operations tasks and limited engineering time. How do you decide which ones to automate, which to leave manual, and which to remove the need for entirely?

level: principalimportance: should knowfreq 42%

basics

~20 s

Rank by payoff against cost: how often the task occurs, how long it takes, and what it risks, weighed against building and maintaining the automation forever. Always ask first whether the task can be eliminated — fixing the cause beats automating the symptom.

open as a page

Your team has measured toil at roughly 70% of engineering time for two quarters running, well over the 50% ceiling you committed to. As the team's lead, what do you actually do about it?

level: principalimportance: should knowfreq 40%

basics

~20 s

Treat the ceiling as a commitment, not a metric. Attack the two or three dominant toil sources, ring-fence capacity so the fix actually happens, and cap intake or hand work back to service owners. Hiring to absorb toil makes it permanent.

open as a page

An operations tool offers a dry-run mode that prints the changes it would make without making them. What does running that mode actually prove before you let the tool loose on production, and what can it still not tell you?

level: juniorimportance: nice to knowfreq 40%

basics

~20 s

A dry run proves the tool's inputs, targeting and scope: which objects it selected and how many it would change. It cannot prove the change succeeds, that permissions allow it, that the world will still look the same at run time, or that a partial failure is safe.

open as a page

A repetitive manual task on your team could clearly be automated. Under what circumstances is leaving it manual the right call, and how would you make that judgement defensible to your team?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Leave it manual when payback fails — build plus maintenance cost exceeds the hours saved over the task's remaining life — or when the action is rare and destructive enough that rarely exercised automation is riskier than a careful human.

open as a page

Leadership mandates that every service must have runbooks covering its top failure modes before launch, and compliance is tracked as a runbook count per service. Would you adopt that, and what would you commit to instead?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

No — counting documents measures writing, not response capability, and produces procedures written from imagination that nobody has executed. Commit instead to coverage defined against the alerts that actually page, at least one execution by someone who did not write it, and a documented escalate-to-owner entry as a legitimate answer.

open as a page