Your team has measured toil at roughly 70% of engineering time for two quarters running, well over the 50% ceiling you committed to. As the team's lead, what do you actually do about it?
answer
- a ceiling with no consequence is decoration
- break it down before you fix it
- two or three sources dominate
- protected capacity, or it never happens
- headcount makes the toil permanent
basics
~20 sTreat the ceiling as a commitment, not a metric. Attack the two or three dominant toil sources, ring-fence capacity so the fix actually happens, and cap intake or hand work back to service owners. Hiring to absorb toil makes it permanent.
solid answer
~50 sA ceiling with no consequence is decoration, so the first move is to make the breach cost something. Practically: break the 70% down and find the two or three sources that dominate — it is almost never evenly spread. Then ring-fence capacity to fix them, because a team at 70% will never find the time incidentally; that usually means an interrupt-catcher rotation shielding the rest of the team, and it costs the shielded work of whoever is on it. In parallel, reduce inflow, which is the lever people avoid: stop onboarding new services until launch-readiness criteria are met, or hand specific classes of work back to the owning teams. Hiring is the tempting answer and it is the trap — adding people to absorb toil converts a temporary overload into permanent headcount and removes the pressure that would have fixed it. And check the diagnosis first: 70% in a two-person team supporting a legacy estate may be a staffing or scope problem rather than an automation one.
go deeper
Know that SRE sets a ceiling on toil — around 50% in Google's model — and that exceeding it is meant to trigger action rather than simply be noted in a report.
Be able to name concrete responses: dedicate capacity to reduction, automate the largest source, or push work back to the service owners, and explain why a team at 70% cannot find the time incidentally.
Show judgement about diagnosis. Distinguish an automation problem from a scope, quality or boundary problem, and be explicit about which lever you pull and what it costs the team politically.
Own the commitment itself: define what the breach obliges the organisation to do, hold the line when intake pressure arrives, and refuse the headcount answer unless it comes with a plan that changes the trajectory.
## Why the ceiling exists The roughly 50% figure from Google's SRE guidance is not arbitrary and it is not a law of nature. It is a forcing function. A team above the line has no capacity left to make next quarter cheaper than this one, so it stays above the line — and worse, it will keep saying yes to new work while getting slower, because the compounding is invisible from inside. The ceiling exists so that crossing it triggers something specific. If nothing happens when you cross it, you do not have a ceiling; you have a chart. Two quarters is the important detail in this scenario. One quarter over could be a bad launch or a couple of nasty incidents. Two consecutive quarters is a structural statement about the team's workload, and it should escalate. ## Diagnose before you act Break the 70% down by source. Toil distributions are heavily skewed: two or three categories usually account for most of it. Attacking the long tail feels productive and moves the number by two points; attacking the top source moves it by fifteen. Insist on that breakdown before anyone proposes a solution. Then ask what kind of problem this actually is, because "more automation" is only one of the possible answers: - **An automation problem** — the work is genuinely automatable and nobody has had the time. The classic case, and the one the ceiling is designed for. - **A scope problem** — the team owns more services than a team that size can carry regardless of tooling. No amount of automation fixes a mismatch between scope and headcount. - **A quality problem** — the toil is generated by a service that is simply unreliable, and the interrupts are symptoms. Here the fix belongs in the product, not in your runbooks, and the conversation is with its owners. - **A boundary problem** — your team has absorbed work that belongs to the service teams, one favour at a time, and never handed it back. The response differs completely across those four, and misdiagnosing costs you a quarter. ## The levers, and what each one costs **Ring-fence capacity.** The only reliable way a 70% team ships a fix is to make the fixing time structurally unavailable for interrupts — typically a dedicated interrupt-catcher on rotation, with everyone else genuinely protected. Cost: the person on rotation ships nothing that week, and someone has to enforce the protection when a director walks over with an urgent request. Without that enforcement this degenerates into everyone doing toil again by Wednesday. **Cap intake.** Stop accepting new services, or gate them behind launch-readiness criteria the owning team has to meet. This is the most effective lever and the most politically expensive one, because it makes your constraint someone else's problem — which is exactly the point. It only works if you can show the measured number and the trend. **Hand work back.** Return specific classes of work, or the pager itself, to the teams that own the services generating it. The strongest version is explicit: SRE support is conditional on the service meeting a reliability bar, and below that bar the developers carry their own operational load. It is the sharpest incentive in the whole model, and it will cost you goodwill in the short term. **Eliminate the source.** Frequently the top toil category traces to a defect nobody has prioritised: the job that wedges weekly, the deploy that needs manual intervention. Fixing the underlying defect removes the work rather than making it cheaper, and it is usually smaller than building the automation would have been. **Shed or sunset.** Some services are not worth their operational cost. Retiring one, or dropping it to a lower support tier, is a legitimate and underused answer. **Hire.** Sometimes correct — if the diagnosis is genuinely scope — but understand what it does. Adding an engineer to absorb toil at 70% means you now have permanent funding for that toil and no pressure to remove it, and toil that scales with the estate will overrun the new headcount too. Hiring buys time; it does not buy a slope change, and it must be paired with a plan that does. ## What a strong answer sounds like It names a breakdown before a remedy, picks two or three levers rather than all of them, and states the cost of each out loud — including who will be annoyed. It resists the two comfortable answers: "we'll automate our way out" (with what time?) and "we need more people" (which makes it permanent). One more distinction worth being crisp about: the toil ceiling is not the error budget. They measure different things — engineering time consumed versus reliability spent against a target — and they trigger different conversations. Conflating them is a common tell that someone has read about both and operated neither.
- Isn't hiring another engineer the obvious fix for a team at 70% toil?It relieves the symptom and entrenches the cause. You now have permanent funding for that toil, the pressure that would have driven the fix is gone, and if the toil scales with the estate the new engineer is consumed within a year or two. Hiring is defensible when the diagnosis is scope rather than automation — but only paired with an explicit plan for what the added capacity will remove.
- Most of the toil turns out to be absorbed by one senior engineer who is good at it. What does that change?It makes it more urgent, not less. Concentration hides the problem from the average, creates a single point of failure, and burns out your most capable operator. Deliberately spread the interrupt load — a rotation everyone takes — so the pain is visible to the whole team and to management. Comfort with the current arrangement is exactly why it has survived two quarters.
- How do you choose which toil source to attack first?By share of the total, then by tractability. Rank the categories, take the largest, and check whether the fix is automation, a defect fix, or handing the work back — the cheapest of those is often not the automation. Attacking anything outside the top few is motion without movement: it feels productive and moves the headline number by a couple of points.
saying these in an interview costs you the question
- Proposes automating everything with no capacity to do it
- Hires to absorb toil with no plan to remove it
- Treats the 50% figure as an inviolable industry rule
- Confuses the toil ceiling with the error budget
- Attacks small toil sources while the dominant one stands
- Reports the breach upward with no proposed lever