skip to content

A remote lookup is wrapped in retry and then in caching — what changes when the two wrappers swap places?

level: middleimportance: must knowfreq 56%

answer

  1. calls enter from the outside in
  2. the outer wrapper can short-circuit the inner
  3. hit returns before retry is entered
  4. retry outside re-enters the cache each attempt
  5. a stored failure makes retrying pointless

basics

~20 s

The outer wrapper runs first and decides what the inner one ever sees. With caching outside, a hit returns without any retry happening. With retry outside, every attempt goes through the cache first, so the cache shapes each attempt.

solid answer

~40 s

Stacked wrappers form an onion, and a call enters from the outside in. With `caching(retry(lookup))` the cache is consulted first: a hit returns immediately and the retry wrapper is never entered, so retries only ever happen on a miss and the stored value is whatever finally succeeded. With `retry(caching(lookup))` the retry wrapper is outermost: each attempt re-enters the caching wrapper, so the cache is consulted once per attempt. If that inner cache only stores successes, the retries still reach the remote call and the ordering mostly wastes lookups; if it also stores failures, every attempt reads the same stored failure and retrying can never succeed. The same reasoning places a timing wrapper: outside the cache it measures what callers feel, inside it measures only real lookups.

code

pseudocode · 11 lines
pseudocode
function caching(inner)
    store = emptyMap()
    return function(key)
        if store.has(key)
            return store.get(key)      // inner is never called
        value = inner(key)             // a raised error skips the next line
        store.put(key, value)
        return value

fast = caching(retry(lookup))   // hit: retry is not entered at all
slow = retry(caching(lookup))   // each attempt re-enters caching first

go deeper

for a junior

Learn to read the nesting out loud: the outermost wrapper is the one the call reaches first, and the innermost is the real function. Being able to name that order is most of the answer.

for a middle

Trace both orderings of retry and caching over a failing call, and say for each what got stored, how many outbound calls happened, and which wrapper was never entered.

for a senior

Show this as an operational story: which metric moved when the cache was introduced, why the attempt counter looks wrong, and where you moved a wrapper so the numbers meant what the dashboard claimed.

for a principal

The lead's angle is standardising the order across services, so the same metric name means the same thing everywhere, and deciding which behaviours are allowed to short-circuit others at all.

## The stack is an onion Wrappers compose by nesting: `caching(retry(lookup))` means the caching wrapper holds the retry wrapper, which holds the raw lookup. A call **enters from the outside in** and a result **leaves from the inside out**. The outermost wrapper therefore gets first refusal on every call — it can answer without ever calling inward — and it is the last to touch the result on the way back. That single fact is why order is not cosmetic. Each wrapper only sees the calls the layer above it decided to pass down. ## Caching outside retry Order: `caching(retry(lookup))`. - A **cache hit** returns the stored value and the retry wrapper is never entered. No attempt is counted, no backoff is waited. - A **miss** falls through to the retry wrapper, which calls the raw lookup and calls it again on failure until it succeeds or gives up. - What gets stored is the value that eventually succeeded, no matter how many attempts it took. The cost of the retries is paid once and amortised over every later hit. - If the retry wrapper gives up and raises, the storing step never runs — a raised error skips it — so failures do not become stored answers. This is usually the order you want: the expensive machinery sits behind the cheap check. ## Retry outside caching Order: `retry(caching(lookup))`. - The retry wrapper is entered first, and **each attempt re-enters the caching wrapper**. - Attempt one misses and calls the raw lookup. If it fails, nothing is stored, so attempt two misses again and calls the raw lookup again. The cache has done no work at all for this call. - Once a value is finally stored, later calls hit on their first attempt — so the arrangement is not broken, just wasteful for the failing path. - **If that inner cache also stores failures**, the arrangement is actively wrong: attempt two reads the stored failure instead of calling the remote service, and every remaining attempt does the same. The retry loop burns its budget without ever making a new request, and can never succeed. | Order | What a cached value does | What a failure does | |---|---|---| | caching outside retry | returns at once; retry never entered | retried inside, then propagates; nothing stored | | retry outside caching | hit on the first attempt ends the loop | each attempt re-enters the cache before calling out | ## Where the timer goes The third wrapper makes the point sharper. Place the timing wrapper **outside** the cache and it measures what callers actually experience, hits included — which is the honest latency number, and also the one that collapses towards zero as the hit rate climbs. Place it **inside**, next to the raw lookup, and it measures only calls that really went out, so it answers a different question: how slow is the dependency. Inside the retry wrapper it times one attempt; outside it, one attempt plus every retry and every wait between them. None of these is wrong; they are four different metrics, and the position is the definition. ## Choosing an order 1. **Put the cheapest short-circuit outermost.** Anything that can answer without calling inward belongs above the machinery it saves you from. 2. **Put anything that must see every real call innermost.** A wrapper that counts outbound requests has to sit below the cache, or it will count answers that never left the process. 3. **Decide what each measurement means and then place the timer,** rather than placing it and interpreting whatever comes out. 4. **Write the order down at the wiring point.** The nesting is the only place this behaviour is specified; a reader cannot recover it from the wrappers themselves. Whether the cache in any of this is sound, and how large it is allowed to grow, is a separate question from where it sits — placement is what the stacking order decides. ## The trap in interviews Candidates often answer that swapping the two is a performance detail. It is not: with one ordering a failure is retried and then cached, and with the other a cached failure prevents the retry from ever happening. Same two wrappers, same lookup, different behaviour — and the only thing that changed was which one was applied last.

  • With caching outermost, what does the retry wrapper's attempt counter actually count?
    Only attempts made on a cache miss. Hits return above it and are never seen, so the counter measures the cost of misses, not of calls. Read as 'attempts per call' it will look far too low; read as 'attempts per miss' it is correct. The position decides which of the two the number means.
  • Where would you place a logging wrapper that must record every request the process actually sends outward?
    Innermost, directly around the raw lookup and below both the cache and the retry wrapper. Above the cache it would log answers that never left the process; above the retry wrapper it would log one line for a call that in fact made several requests.
  • Does swapping two wrappers ever leave behaviour identical?
    Yes, when neither can short-circuit the other and neither changes what the other sees — two independent observers, such as one counting calls and one logging arguments, commute. Order matters exactly when a wrapper can answer without calling inward, call inward more than once, or alter what passes through.

saying these in an interview costs you the question

  • Says wrapper order is only a performance detail, not a behaviour difference
  • Believes a cache hit still passes through the retry wrapper underneath it
  • Assumes a timing wrapper measures the same thing wherever it is placed
  • Puts an outbound-request counter above the cache and reports it as traffic
  • Cannot say which wrapper a call reaches first in a given nesting