skip to content

How do you choose which metric to quote when a backend change improved several at once?

level: middleimportance: should knowfreq 71%

answer

  1. Closest to what your change touched
  2. Money impresses, minutes are defensible
  3. Every metric needs a named source
  4. Activity counts are not impact
  5. One claim, the rest as follow-up

basics

~20 s

Quote the measure closest to the pain your change removed, in units your listener already understands, and only if you can say where the value came from. Bigger downstream measures such as revenue impress more but are far harder to defend.

solid answer

~40 s

Three filters, applied in order. First, **proximity**: pick the measure your change acted on directly — for a batching rewrite in queue consumers that is drain time, not the quarter's revenue. Second, **sourcing**: if you cannot name where the before-value and after-value came from, the metric is not usable, however impressive. Third, **legibility**: the reader should not need your internal glossary, so either use a common measure or define it in the same breath. When several improved, lead with one and let the rest be follow-up material — `drain time from 41 minutes to 9 across 3 of 17 consumers` is one clean claim, and the reduced page volume and infrastructure saving are things you offer when asked.

go deeper

for a junior

Learn the metric families — speed, reliability, cost, adoption, business outcome — and practise naming which one your last piece of work actually moved. Avoid quoting tickets closed or code volume as a result.

for a middle

Explain why a measure sitting close to your change is easier to defend than a downstream one, and be able to name the source you would quote for each end of the number.

for a senior

Demonstrate the judgment to lead with one defended figure and hold the rest back. Show that you know which of your metrics would survive an interviewer walking the causal chain backwards.

for a principal

Own the choice of what your work is measured by in the first place. Be ready to argue why the measure you optimised for was the right one, including what it deliberately left unimproved.

## The families of metric available to an engineer Most backend work moves something in one of five families, and they differ enormously in how easy they are to defend. | Family | Typical measures | Attribution to you | |---|---|---| | Latency / throughput | queue-drain time, request latency, jobs processed per hour | strong — the change acts on it directly | | Reliability | pages per week, failed runs, error rate, incident count | strong, but noisy over short windows | | Cost | compute or storage spend, instance count, per-request cost | good, if you can isolate your workload from the bill | | Adoption | teams onboarded, services migrated, internal users | good, though it measures uptake rather than benefit | | Revenue / business outcome | conversion, retention, orders completed | weak for an individual engineer — long causal chain | The pull is always toward the bottom of that table, because money sounds more serious than minutes. The evidence runs the other way. A behavioral interviewer weighing a revenue claim from a queue change has to accept a chain with several links they cannot see, and the honest ones will simply ask you to walk the chain. If you cannot, the claim converts into a doubt about everything else on the page. ## Filter one: proximity to the change Ask what your change physically did, then name the measure that sits closest to it. Rewriting batching in three queue consumers acts on how fast a backlog drains. Drain time falling from 41 minutes to 9 is a direct consequence you can explain mechanically — fewer round trips per batch, less contention on the write path. Anything further downstream is a consequence of a consequence, and every extra link is a place where somebody else's work could explain the movement instead of yours. This does not mean the downstream effect is off limits. It means you state the near measure as the claim and offer the far one as context: the drain time is the number, and the fact that the nightly settlement finished before the support shift started is the reason anyone cared. ## Filter two: can you source both values? A metric is only usable if you can answer where each end of it came from. Alert timestamps, a dashboard panel you looked at weekly, the job's own run log, an incident timeline, a cost report — any of these is a source. Memory is not. If you find yourself unable to name the source for the before-value, that is a signal to pick a different measure rather than to quote the one you like best and hope the question does not come. ## Filter three: legibility to this listener Internal measures do not travel between employers. A score, an index, or a team-specific ratio means nothing outside the company that invented it, and spending your answer defining it burns the time you wanted for the mechanism. Prefer measures with self-evident units — minutes, count of services, requests per second, currency of spend — or define an internal one in a single clause: *the backlog age we alerted on, measured at the oldest unprocessed message*. ## What to leave off Activity counts are the most common substitute for impact: tickets closed, pull requests merged, services touched, lines of code. They measure motion, not consequence, and experienced interviewers read them as an admission that nothing measurable moved. The same applies to size-of-system boasts — a fleet of seventeen services is context for your work, not a result of it. ## When several measures moved at once Pick one to be the claim. A bullet or a closing sentence carrying three numbers reads as a dashboard, and the reader retains none of them. Lead with the one that best satisfies all three filters, and hold the others in reserve for follow-ups, where they land as depth rather than clutter. If the secondary effects are the more interesting story — the on-call rotation stopped being woken by that backlog — say that in words, and keep the single defended figure as the anchor. ## The tell that you chose badly If your metric requires you to explain why it is impressive before you can quote it, you chose the wrong one. Good metrics land in one sentence and immediately invite a mechanical question. Bad ones require a preamble about how the company measures things, and the interviewer's attention is spent by the time you reach the value.

  • Would you ever put a revenue figure on a bullet about a queue-batching change?
    Only if I could walk the chain from my change to the money without hand-waving, and normally I cannot. What I do instead is state the operational number I own — drain time from 41 minutes to 9 — and describe the business consequence qualitatively, such as the settlement run finishing before the support shift began. That keeps the defended claim and the motivation separate, which is where they belong.
  • What if the only measure that moved is one I would not recognise?
    Then define it inside the sentence, in a clause, not a paragraph: the backlog age we alerted on, measured at the oldest unprocessed message. If it cannot be defined that briefly, look for a proxy with obvious units — minutes, failed runs, spend — and mention the internal measure only if you are asked to go deeper.

saying these in an interview costs you the question

  • Picks the largest-sounding metric rather than the one they can source
  • Quotes activity counts such as tickets closed as impact
  • Claims a revenue effect with no chain from the change to the money
  • Uses an internal score without saying what it measures
  • Stacks three or four figures into a single unmemorable claim

context