In SRE, what makes operational work count as "toil", and name two kinds of manual work that are not toil?
answer
- not every manual task qualifies
- six tests, roughly all must hold
- nothing enduring left behind
- grows as the service grows
- overhead and engineering are different buckets
basics
~20 sToil is manual, repetitive, automatable, tactical work with no enduring value that scales with the service. Manual work failing those tests — a one-off migration, or novel debugging of a new failure — is engineering or overhead, not toil.
solid answer
~50 sToil is a specific category, not a synonym for unglamorous work. The criteria are: it is **manual**, **repetitive**, **automatable**, **tactical** (reactive, interrupt-driven), it produces **no enduring value** — the service is in the same state afterwards as before — and it **scales linearly with service growth**. A task has to hit essentially all of those to count. So a one-off hardware migration is manual but not repetitive and leaves lasting value: that is project work. Debugging a novel failure mode is manual and reactive but not automatable, because you don't yet know what to encode: that is engineering. Team meetings, interviews and expense reports are overhead, not toil. The reason SRE bothers with a precise definition is budgeting: toil is the number you point at to justify spending engineering time on automation, and Google caps it at roughly 50% of an SRE's time. A definition loose enough to cover everything you dislike can't carry that argument.
go deeper
Be able to state the definition in one sentence and list the criteria, and say plainly that toil is necessary work, not bad work — it is the share of time it consumes that is the problem.
Explain why each criterion is there, especially automatable and no-enduring-value, and be ready to classify a borderline item such as a quarterly certificate renewal or a customer data-export request.
Show that you use the definition as an instrument: point at a real ticket queue, say which fraction of it is toil, and defend the calls to a manager who wants the number smaller.
Own the boundary across teams so the number means the same thing everywhere. A definition that drifts per team cannot be aggregated, and a loose one gets used to justify whatever automation someone already wanted to build.
## Why the word needs a definition at all Every operations team has work it resents, and every team would like that work automated. "Toil" is SRE's attempt to make that argument fundable rather than emotional. If toil is precisely defined and measured, a team can say "41% of our engineering time went to toil last quarter, here are the three sources, here is what we intend to spend to remove them" — and a manager can act on it. If toil just means "work we don't enjoy", the number is unfalsifiable and nobody will fund anything on the strength of it. ## The criteria Google's Site Reliability Engineering book gives six tests. Work is toil when it is: - **Manual** — a human has to touch it. Time spent hand-running a script counts; the script running itself does not. - **Repetitive** — you have done it before and you will do it again. First time is investigation; the twentieth time is toil. - **Automatable** — a machine could plausibly do it. If the task requires judgement no one has yet been able to encode, it is not toil, it is engineering. This is the criterion that keeps incident diagnosis out of the bucket. - **Tactical** — it is reactive and interrupt-driven rather than strategy-driven. A page you answer at 3am is tactical; the capacity plan you write on Tuesday is not. - **Devoid of enduring value** — when you are done, the service is in the same state it was in before. You restarted the stuck worker; nothing about tomorrow got better. - **Growing at least linearly with service size** — the cost of the task rises as the estate grows. This is what makes toil an existential problem rather than an annoyance: work that scales with the service eventually consumes the whole team. In practice you rarely need all six to make a call, but the two that do the most classifying work are *automatable* and *no enduring value*. If a task leaves something permanently better, or if nobody knows how a machine would do it, it is not toil. ## What is not toil Three categories get misfiled constantly. **Overhead.** Meetings, performance reviews, hiring loops, expense reports, planning. These are manual, repetitive and arguably valueless, but they are not operational work on your service, they do not scale with it, and your team cannot automate them away. They belong in a separate line of the time budget. Folding them into the toil number inflates it and makes the whole measurement easy to dismiss. **Project and engineering work.** A migration, a schema change, a new deployment pipeline. Manual, sometimes deeply tedious, but each one leaves the system permanently different. Doing it once is investment, not toil. If you find yourself doing the "one-off" migration for the eleventh time, it has become toil and you should say so. **Genuine incident response.** The novel outage where you don't know what is broken is the opposite of toil: it is not repetitive, not automatable, and the understanding it produces is enduring value. What *is* toil is the incident you have seen fifteen times, where you follow the same steps to the same resolution. ## Toil is not bad work, and it is not zero Two corollaries candidates often miss. First, toil is legitimate, necessary work. Someone genuinely does have to restart the worker. Calling it toil is not calling it beneath anyone, and an SRE who refuses to do operational work has misread the idea entirely. The claim is only that its *share* of the team's time has to be bounded, because a team at 100% toil has no capacity left to make next quarter cheaper than this one. Second, the target is not zero. Google's published guidance is a ceiling of roughly 50% of an SRE's time, precisely because some toil is the price of running a real service and because keeping engineers in touch with production has value. A team reporting 0% toil is either not operating anything or not measuring. ## Applying it The useful exercise in an interview is classification. Given a queue — quarterly certificate renewals, a customer's data-export request, a flaky test rerun, a capacity review, a manual production release, a novel latency investigation — say which are toil and why, and the honest answer will not be "all of them". Certificate renewal and the manual release are textbook toil: repetitive, automatable, leaving nothing behind, and each new service adds another one. The capacity review and the latency investigation are engineering. The export request is toil only if it keeps coming; if it is the first, it is a signal you may need a self-service feature. The classification is the whole point: it converts a vague complaint into a prioritised list of things worth building.
- Is toil the same thing as operational work?No. Operational work includes incident response to novel failures, capacity reviews and design consultation — none of which are automatable or repetitive. Toil is the automatable, no-enduring-value subset. Teams that equate the two end up concluding that all ops work is bad, which is both wrong and corrosive to the relationship with the service owners.
- If a task is manual and repetitive but takes five minutes a month, is it toil?Yes, by definition — but that says nothing about whether to automate it. The definition classifies; the measurement prioritises. Five minutes a month is a rounding error and probably stays manual forever. Being toil and being worth removing are separate judgements, and conflating them is how teams end up automating trivia while the real cost sits untouched.
- Where do meetings, interviews and expense reports fit?Overhead, tracked as its own line. They fail the criteria on two counts: your team cannot automate them, and they do not scale with the service. Folding them into the toil number inflates it, and the first manager who spots that will discount the whole measurement — which costs you the argument you were trying to fund.
saying these in an interview costs you the question
- Calls every manual or unglamorous task toil
- Treats toil as work beneath senior engineers
- Says the goal is zero toil
- Counts meetings and admin as toil
- Classifies novel incident debugging as toil
- Thinks toil means the work is optional or unimportant