skip to content

A timing interceptor sits inside the retrying interceptor in one chain and outside it in another - what differs in what each records?

level: seniorimportance: must knowfreq 55%

answer

  1. position decides entries and interval
  2. a retrying handler calls inward repeatedly
  3. inner handlers run once per attempt
  4. waits sit in the retrying handler's frame
  5. the gap between the timers is retry cost

basics

~20 s

Position changes both how often a handler runs and what interval it covers. Inside the retrying handler, timing is entered once per attempt and records each attempt alone. Outside it, timing is entered once and records every attempt plus the waits between them.

solid answer

~40 s

A retrying handler calls inward more than once, so everything nested inside it is entered once per attempt while everything outside it is entered once per request. The inner timer therefore emits one measurement per attempt, each covering a single try and excluding the waits between tries; the outer timer emits one measurement per request, covering all attempts, all backoff waits, and the retrying handler's own overhead. Neither is wrong - they answer different questions. The inner one tells you how slow a single call to the target is; the outer one tells you what the caller actually waited. A dashboard built on the inner one can look healthy while every caller is timing out, because the retries and the waits are invisible to it.

code

pseudocode · 9 lines
pseudocode
// A: timing nested inside the retrying handler
chain = wrap(retryHandler, wrap(timingHandler, target))
// 3 attempts -> timingHandler entered 3 times
// each sample = one attempt, no waiting time

// B: timing outside the retrying handler
chain = wrap(timingHandler, wrap(retryHandler, target))
// 3 attempts -> timingHandler entered once
// the one sample = attempt1 + wait + attempt2 + wait + attempt3

go deeper

for a junior

Recall that a handler which retries calls inward more than once, so anything nested inside it runs again on every attempt.

for a middle

Explain both effects of position: how many times the handler is entered, and how much of the elapsed time its open frame spans.

for a senior

Show the operational consequence - a dashboard that looks healthy while callers time out - and say which placement answers which question about the system.

for a principal

Decide what the platform's default placement is and why, so that every team's latency numbers mean the same thing across services.

## Position decides two independent things When handlers nest around one call, a handler's position decides: 1. **How many times it is entered.** Every handler nested inside a handler that proceeds more than once is entered once per inward call, not once per request. 2. **What interval its after-work covers.** A handler's frame stays open across everything nested inside it, so its measurement spans that whole subtree - including time that is not spent in the target at all. A retrying handler is the clearest case because it does both: it calls inward repeatedly, and it spends time between those calls waiting. ## Timing nested inside the retrying handler Here the order is retry, then timing, then the target. The retrying handler makes attempt one; the timer starts and stops around that attempt alone. If the attempt fails, the retrying handler waits, then calls inward again, and the timer runs a second time. With three attempts the timer is entered three times and emits three measurements. What that gives you: - one sample per attempt, so failed attempts are measured as well as the successful one; - no waiting time in any sample, because the wait happens in the retrying handler's own frame, outside the timer; - a count of measurements that no longer equals the number of requests, which quietly breaks any rate derived from it. ## Timing outside the retrying handler Here the order is timing, then retry, then the target. The timer is entered once. Its frame stays open for every attempt, every wait between attempts, and the retrying handler's own bookkeeping, and it stops only when a result finally travels outward. | question you are asking | put the timer inside the retrying handler | put the timer outside it | |---|---|---| | how often is it entered | once per attempt | once per request | | does a sample include backoff waits | no | yes | | does a failed attempt produce a sample | yes, one per attempt | no, the request produces one sample | | what a high percentile means | one slow call to the target | one slow experience for the caller | | what it hides | that the request was retried at all | which attempt was slow, and how many there were | ## The same rule bites elsewhere in the same chain - **Quota inside a retrying handler is charged per attempt**; outside it, per request. One logical request can therefore consume several units of a budget that was sized per request. - **A handler that can stop the call before the target runs makes everything nested inside it optional**, so a counter placed inside such a handler undercounts requests by exactly the number that were stopped. - **A handler that rewrites arguments is only visible to handlers nested inside it**; move it inward and the handlers that used to see the rewritten form now see the original. - **A failure translated by a handler looks like a different failure to everything outside it**, so a classifier's position decides which name the outer handlers log. ## Choosing, rather than inheriting, the position The useful habit is to decide, for each handler, what question its numbers must answer and then place it so that its open frame spans exactly that. Caller-facing latency, availability as the caller experiences it, and end-to-end budgets belong outside the retrying handler. Target health, per-attempt cost, and anything you intend to compare against the target's own instrumentation belong inside it. Many systems keep both, precisely because the gap between the two is the retry cost, and that gap is the number that tells you whether retries are helping or amplifying load. What makes this a senior question is that the wrong choice produces a system that is *quietly* wrong rather than broken: both chains work, both emit plausible numbers, and the difference only surfaces during an incident, when the dashboard says the target is fast and every caller says the service is slow. The candidate who can explain the mechanism - one frame per handler, repeated inward calls for everything nested inside - can predict that mismatch from the assembly order alone, without measuring anything.

  • A retrying handler makes three attempts. How many times is a counter nested inside it incremented, and how many times is one outside it incremented?
    Three times inside, once outside. Everything nested inside a handler that proceeds repeatedly is entered once per inward call. That is why per-request budgets should be enforced outside the retrying handler and per-attempt load accounted inside it.
  • Both timers are in the chain at once. What does the difference between them tell you?
    The retry overhead: waiting time plus the cost of the attempts that failed. If that gap grows while the inner measurement stays flat, the target is failing more often rather than getting slower, and the retries are adding load instead of absorbing a blip.
  • Where should a handler that enforces an overall deadline for the request sit?
    Outside the retrying handler, because the deadline covers the whole request including every retry and wait. Nested inside, it would restart on each attempt and permit a total elapsed time of roughly attempts multiplied by the per-attempt limit.

saying these in an interview costs you the question

  • Assumes every handler is entered exactly once per request
  • Thinks a handler nested inside a retrying handler sees the waits between attempts
  • Says the two placements differ only in overhead, not in meaning
  • Believes a percentile from the inner timer reflects caller-visible latency
  • Treats handler order as a cosmetic registration detail