skip to content

Two recurring manual tasks each cost your team about an hour a week today. One's frequency grows in proportion to the number of services you run; the other is fixed no matter how large the estate gets. Why does SRE's definition of toil treat those differently, and which do you attack first?

level: middleimportance: should knowfreq 48%

answer

  1. same cost today, different futures
  2. think O(1) versus O(n) in human hours
  3. the slope, not the hour
  4. check what the multiplier actually is
  5. cheap wins are not a priority argument

basics

~20 s

Toil that scales with the service is the dangerous kind: its cost rises with the estate until it consumes the team. Attack the scaling task first — the fixed hour stays one hour forever, however irritating it is.

solid answer

~50 s

They cost the same today, but they have different futures, and that is exactly why "grows at least linearly with service growth" is one of the criteria for toil. The fixed task is a constant: an hour a week at ten services, an hour a week at a hundred. The scaling task is a function of the estate. If you are adding services at, say, 50% a year, an hour a week becomes about ninety minutes next year, three and a half hours in four years — and at that point a meaningful slice of an engineer exists only to keep doing it. That is the work that eventually caps how much your team can run. So I would put the structural fix — self-service, a platform default, or eliminating the need — on the scaling task, and treat the fixed one as opportunistic: worth a day's scripting if someone has a day, never worth displacing the scaling fix.

go deeper

for a junior

Know that one of toil's defining criteria is that it grows with the service, and be able to say why work that grows is treated as more urgent than work that stays constant.

for a middle

Do the arithmetic out loud: project the scaling task forward at a stated growth rate, show when it becomes a fraction of a person, and use that to justify the ordering.

for a senior

Show that you check the growth driver empirically rather than assuming it, and that you can defend spending a month on the structural fix while the cheap win sits undone.

for a principal

Own the sublinear-scaling goal across the org: the question is not which task to script but which classes of per-service work the platform should absorb so team size stops tracking service count.

## The criterion that does the real work Of the criteria that define toil, "grows at least linearly with service size" is the one that turns an annoyance into an argument. Manual, repetitive and low-value describe how the work *feels*. Scaling with the service describes what it *does to you over time*, and that is what a manager can be shown on a chart. The distinction is the same one you would draw about algorithmic complexity, applied to human time. Some operational cost is O(1) in the size of the estate: the weekly ops review, the monthly key rotation for one shared credential, the quarterly audit form. Some is O(n): a certificate renewal per service, an access grant per customer, a manual release per deployable, a capacity check per cluster. Both cost an hour today. Only one of them has a slope. ## Doing the arithmetic out loud The reason to carry the numbers is that they change the decision. Suppose the O(n) task is one hour per week across the estate you have now, and the estate grows 50% a year: ``` year 0: 1.0 h/week (~52 h/year) year 1: 1.5 h/week year 2: 2.3 h/week year 3: 3.4 h/week year 4: 5.1 h/week (~265 h/year, ~0.13 FTE) ``` At a doubling rate it is far worse — an hour a week becomes sixteen hours a week in four years, which is not "a nuisance", it is a hiring request. Meanwhile the fixed task has cost exactly 52 hours a year the entire time. And this is one task. A team typically carries several O(n) items at once, and they compound, which is why teams that never fix the slope find that adding services stops being possible long before the infrastructure runs out. The tell in a real organisation is a launch queue: new services waiting because operations cannot absorb them. ## The decision, with its cost Attacking the scaling task first is the right default, but state the cost honestly, because the fixed task is usually the cheaper win. Structural fixes to O(n) toil are rarely a script. Per-service certificate renewal is fixed by automated issuance and rotation being the platform default, not by a better renewal runbook. Per-customer access grants are fixed by self-service with an approval path, not by a faster ticket template. That is weeks of work touching systems other people own, with a negotiation attached. The fixed task, by contrast, is often an afternoon. So the practical answer is about *priority*, not sequence. Take the cheap fixed win if it is genuinely an afternoon and it removes an interrupt — interrupt reduction has value beyond the hours. But do not let a queue of cheap wins substitute for the structural fix, because that is the failure mode: a team that spends two years automating easy things while the slope quietly eats it. ## When the scaling task is not the priority Three honest exceptions, and being able to name them is what separates a considered answer from a recited rule. **The growth curve is flat over your horizon.** If the estate is not growing — a mature product, a system being sunset in a year — then O(n) with a constant n is just O(1), and the slope argument evaporates. Check the actual growth rate rather than assuming it. **The scaling factor is not service count.** Sometimes the multiplier is customers, regions, tenants or data volume, and those grow at completely different rates. "Scales with growth" means growth of the thing that drives the task, and getting that wrong inverts the ranking. **The fixed task is a risk, not just an hour.** A once-a-quarter manual production change that nobody remembers how to do is worth automating or documenting on risk grounds even though its time cost never grows. Time is not the only axis. ## What the interviewer is checking That you distinguish the cost of work today from its trajectory, and that you can defend a prioritisation with an estimate rather than a feeling. The weak answer is "automate both, obviously" — true and useless, because the team has capacity for one of them this quarter. The strong answer names the growth rate, does the multiplication, picks the one with the slope, and says out loud what it costs to go after the harder target first.

  • What does it mean in practice for a team's operational cost to scale sublinearly with the service?
    That each additional service, tenant or region adds less operational work than the last — because the work happens once at the platform level rather than per instance. Automated certificate issuance, self-service provisioning and a shared release pipeline all convert per-service work into fixed work. Sublinear scaling is the actual goal of toil reduction; hours saved is just the proxy you can measure.
  • The fixed task is a day's scripting, the scaling one needs a month of platform work. Does that change your answer?
    It changes the sequencing, not the priority. Take the day if it is really a day and it removes a recurring interrupt. But the month has to get scheduled, not deferred behind an endless queue of cheap wins — that deferral is the most common way teams stay busy and still get overwhelmed. I would commit the platform work to the quarter's plan before spending the day.
  • How would you find out whether a task actually scales with growth?
    Count occurrences per period against whatever drives it — services, tenants, regions, deploys — over the last several months, and see whether the ratio is flat. If occurrences per service are constant and service count is rising, it scales. It is worth checking rather than assuming: tasks people are certain scale sometimes turn out to be driven by one noisy customer instead.

saying these in an interview costs you the question

  • Ranks toil purely by hours spent today
  • Says automate everything without prioritising
  • Assumes the estate always grows
  • Confuses the driver of growth with service count
  • Treats a stack of cheap automations as a strategy

context