skip to content

Your manager asks what fraction of your team's time goes to toil. How do you actually measure that, and how do you keep the number honest enough to make decisions on?

level: seniorimportance: should knowfreq 55%

answer

  1. a fraction of team time, over a quarter
  2. the ticket queue is the meter
  3. unticketed work is the big bias
  4. measure the team, watch the distribution
  5. a target corrupts the metric

basics

~20 s

Measure toil as a share of team engineering time per quarter, from ticketed interrupts, on-call logs and periodic sampling. Keep it honest by ticketing every interrupt, counting per team, and trusting the trend over the exact figure.

solid answer

~50 s

The unit is a percentage of team engineering time over a window long enough to smooth out a bad week — a quarter is typical. The denominator is engineer-hours actually available for work; the numerator comes from three sources. First, the ticket queue: every interrupt gets a ticket with a toil label, even the ones resolved in five minutes, because unticketed work is the single biggest source of undercounting. Second, the on-call record: shifts, incidents and their handling time. Third, a periodic sample — a week each quarter where people log what they actually did — to catch work no queue sees. Measure per team, not per person, or you get defensiveness instead of data. And treat the number as directional: precision beyond a few percentage points is fiction, but a trend from 30% to 55% across two quarters is a decision you can act on.

go deeper

for a junior

Know that toil is reported as a percentage of team engineering time over a period, and that the ticket queue is where the raw data normally comes from.

for a middle

Explain the mechanics: what goes in the numerator and denominator, why a quarter rather than a week, and why unticketed five-minute fixes are the main reason the figure comes out too low.

for a senior

Demonstrate that you have built one of these. Name the biases, keep the collection cheap, break the number down by source, and say which decision each threshold triggers before anyone starts collecting.

for a principal

Own the definition and the incentives across teams so figures are comparable and nobody is graded on theirs. The moment toil percentage becomes a target it stops being a measurement and starts being a negotiation.

## What you are computing Toil is reported as a fraction: toil hours over total engineering hours, for a team, over a window. Both halves need care. The denominator is not headcount times 40. It is the hours actually available for engineering after holidays, leave, training and the overhead that is not toil. Using an idealised denominator systematically understates toil, which is the opposite of the mistake you want to make when you are arguing for investment. The window matters as much as the arithmetic. A week is meaningless — one bad incident swamps it. A year is too slow to steer with. A quarter is the usual compromise, aligned to whatever planning cycle the team already lives in, because the decision the number feeds is a planning decision. ## Where the numbers come from **The ticket queue is the primary meter.** The discipline that makes it work is uncomfortable but simple: every interrupt gets a ticket, including — especially — the ones an engineer fixes in three minutes without being asked. That is the work that never appears anywhere and is the reason measured toil is almost always lower than real toil. The cost is real: you are asking people to file a ticket for something they just did. You buy it back by making filing trivial and by visibly using the data. **On-call records** give you shift counts, alert volumes and incident handling time. Careful here: on-call is only one channel. A team that measures toil purely from its pager will conclude toil is low while the same engineers spend every weekday afternoon on manual access requests. **Periodic sampling** catches the rest. One week a quarter where everyone logs their day into a handful of buckets — toil, project, overhead, incident — is enough. It is self-reported and therefore biased, but you are using it to find categories the ticket queue is blind to, not to produce a precise figure. A worked example makes the shape concrete: ``` Team: 6 engineers x 4 weeks x 40 h = 960 engineer-hours interrupt tickets: 180 x 0.5 h = 90 h on-call handling (logged, 2 rotations) = 60 h manual releases + access requests = 210 h ------------------------------------------------ toil = 360 h = 37.5% ``` The useful move is then to break the numerator down by source, because it is almost never evenly spread. Two or three sources usually dominate, and that list — not the headline percentage — is what you take into planning. ## How the number gets dishonest Four failure modes, and a senior answer names them before the interviewer does. **Invisible work.** The five-minute fix nobody tickets. The engineer who quietly restarts the service every morning before standup. This is the dominant bias and it always points the same way: measured toil is a floor, not an estimate. **Classification drift.** Without an agreed boundary, one engineer books manual releases as toil and another books them as project work. The fix is a short written rule and a couple of worked examples, revisited when a genuinely ambiguous case appears. Numbers that cannot be compared across teams cannot be aggregated, and the first thing a director does is aggregate. **Gaming, in both directions.** A team that wants automation funded has an incentive to inflate; a team whose performance is judged on the number has an incentive to deflate. The moment toil percentage becomes a performance target rather than a planning input, it stops measuring anything. Keep it a team-level planning instrument. **Per-person accounting.** Measuring individuals produces defensiveness and hides the real pattern, which is usually concentration: one engineer absorbing most of the interrupts while the reported team average looks tolerable. Measure the team, but do look at the distribution — a team at 40% where one person is at 80% has a problem the average conceals. ## What the measurement is for A number nobody acts on is itself toil. The measurement exists to feed decisions: whether the team can accept another service, whether this quarter's project slate is realistic, which specific source of toil gets engineering time, whether work has to be handed back to service owners. Before you build the measurement, agree what threshold triggers what response — otherwise you will have a chart and no consequence. And keep the measurement cheap. If tracking toil costs more than a few percent of the team's time, you have added a new source of it and the practice will be abandoned within two quarters, which is precisely how most toil dashboards die.

  • How do you catch toil that never gets a ticket?
    Make ticketing every interrupt the team norm, including instant fixes, and back it with a periodic sampling week where people log their day into a few buckets. You can also mine indirect traces — chat requests, ad-hoc command history on production, approval workflows — to find categories nobody files against. Assume the gap is large: measured toil should always be presented as a floor.
  • Should toil be measured per engineer or per team?
    Per team for the headline number, because individual accounting invites defensiveness and turns a planning input into a performance judgement. But look at the distribution, because concentration is the common failure: a team averaging 40% may be one person at 80% absorbing every interrupt, and that person burns out while the average looks fine.
  • What granularity of accuracy is actually needed?
    Directional. Distinguishing 34% from 37% is fiction given the measurement error, and chasing that precision costs more than it returns. What you need is a category breakdown accurate enough to rank the top few sources, and a trend stable enough that a move from 30% to 55% across two quarters is unambiguous. Decisions hang on the ranking and the slope, not the decimal.

saying these in an interview costs you the question

  • Estimates toil from memory in a planning meeting
  • Measures only on-call pages as toil
  • Uses toil percentage as a performance target
  • Builds a tracking process costing more than it saves
  • Reports a single number with no breakdown by source

context