skip to content

How do you use unittest.mock side_effect to test a retry path that times out twice then succeeds?

level: seniorimportance: should knowfreq 50%

answer

  1. Do not reproduce randomness in a test
  2. Turn a flaky sequence into a script
  3. Exception entries raise, plain entries return
  4. Length must match the expected call count
  5. Exhaustion hides inside a broad except

basics

~20 s

Give the mocked call an iterable side_effect whose first entries are exception instances and whose last entry is the success value: two TimeoutError instances, then the payload. Each call consumes one entry, so the retry runs deterministically.

solid answer

~40 s

Turn the flakiness into a script instead of leaving it to chance. For a feature-flag client that times out intermittently around its 1,200-request-per-minute peak, set `client.evaluate.side_effect = [TimeoutError(), TimeoutError(), {"enabled": True}]`: exception instances inside an iterable are raised and plain entries are returned, so the test drives exactly two failures and one success with no real timing involved. Patch whatever the backoff sleeps on so the test costs nothing. Use the callable form instead when the outcome must depend on the arguments or on accumulated state — failing only the first request per flag key, say. Keep the iterable exactly as long as the expected number of calls: one call too many raises `StopIteration`, and because that is an `Exception`, a broad `except` inside the retry loop will swallow it and spin.

code

python · 20 lines
python
from unittest.mock import Mock


def evaluate_with_retry(client, key, attempts=3):
    last = None
    for _ in range(attempts):
        try:
            return client.evaluate(key)
        except TimeoutError as exc:
            last = exc
    raise last


client = Mock()
client.evaluate.side_effect = [
    TimeoutError("peak load"),
    TimeoutError("peak load"),
    {"enabled": True},
]
print(evaluate_with_retry(client, "new-checkout"))

go deeper

for a junior

Know that a list assigned to side_effect gives one result per call, and that exception instances placed in that list are raised rather than returned.

for a middle

Explain how to order the entries for a fail-then-recover history, why the list length must match the expected number of calls, and why exception instances beat bare classes when the handler inspects the error.

for a senior

Show that you have debugged this: exhaustion raises StopIteration, a broad except swallows it, backoff sleeps must be replaced rather than waited on, and the give-up branch needs its own scripted run.

for a principal

Own the tradeoff between positional scripts and behavioural fakes. A script encodes an implicit call-count expectation that breaks on refactoring; decide when a shared in-memory fake of the dependency is the cheaper long-term asset.

Intermittent failures are the hardest thing to test honestly, because the property you want — "it recovers" — depends on a sequence of outcomes rather than on any single one. `side_effect` with an iterable is the standard Python answer: it converts a probabilistic fault into a deterministic script. ### Script the sequence, do not reproduce the randomness Consider a feature-flag service client that times out intermittently when the service is near its 1,200-request-per-minute peak, and application code that retries a small number of times before falling back to a cached default. The naive instinct is to make the double fail "sometimes". Never do that: a test that fails one run in fifty is worse than no test, because the team learns to ignore it. Instead state the exact history you want to exercise: ```python client.evaluate.side_effect = [ TimeoutError("peak load"), TimeoutError("peak load"), {"enabled": True}, ] ``` Each call consumes one entry. Entries that are exception classes or instances are raised; everything else is returned. So the first two calls raise and the third returns the payload — precisely the two-failures-then-recovery history, every run, in microseconds. ### Prefer instances when the handler inspects the error An exception *class* in the iterable is raised instantiated with no arguments. If the retry logic reads a message, an `errno`, a status or a `retry_after` off the exception before deciding whether the failure is retryable, supply a fully-constructed instance so that branch is actually exercised. A test that recovers from a bare `TimeoutError()` while production reads `exc.response.status` proves less than it appears to. ### Make backoff free A retry loop almost always sleeps between attempts. Leaving that sleep in place makes the test slow and, with jitter, non-deterministic. Replace the specific sleep function the module under test calls with a mock, so the delay is recorded but not taken; the recorded delays are often worth checking too, since a broken backoff calculation is a real production bug. What matters is that no wall-clock time passes. ### Get the length exactly right The iterable neither repeats nor extends. One call past the end raises `StopIteration`, and two failure modes follow from that. First, `StopIteration` is a subclass of `Exception`. A retry loop written as `except Exception: continue` — depressingly common — will swallow the exhaustion signal and keep calling, so an under-length script produces a hanging test or a bogus "all attempts failed" result rather than an obvious error. Second, since PEP 479 (Python 3.7) a `StopIteration` that escapes a generator body is converted into `RuntimeError`, so the same mistake inside generator-based code surfaces as an unrelated-looking crash. When a mock-driven test fails with a mysterious `StopIteration` or `RuntimeError`, the script being one entry short is the first thing to check. The corollary is that the script's length encodes an expectation about how many times the collaborator is called — an implicit assertion. Some teams like that; others find it too tight, because an internal refactor that adds one legitimate call breaks unrelated tests. A pragmatic middle ground is a script long enough for the intended path plus a callable tail, or a callable `side_effect` that never exhausts. ### When the callable form is better An iterable is positional: it says "the third call succeeds", not "calls for this key succeed". As soon as the code under test may reorder calls, batch them, or interleave several keys, positional scripting becomes fragile. A callable that closes over a set of already-seen keys expresses the real rule — "each key times out once, then succeeds" — and is immune to reordering. A callable is also the way to model a circuit that opens after N failures, or a stateful store that returns whatever was last written. ### Keep the test about your code The point of the exercise is the behaviour of the retry policy: how many attempts, what backoff, which exceptions are considered retryable, what happens when every attempt fails, and whether a fallback value is served. All of those are properties of your code, and a scripted `side_effect` is the cheapest way to reach them. Also script the exhaustion case — an iterable of three `TimeoutError` instances with no success at the end — because the "we gave up" branch is the one that runs during a real incident, and it is the branch most often left uncovered. These mechanics are unchanged across Python 3.10–3.14.

  • The retry loop catches broad exceptions. Why can an over-short side_effect iterable hang the test?
    Exhaustion is signalled with `StopIteration`, which subclasses `Exception`. A loop written as `except Exception: continue` swallows it and calls the mock again, which raises again, so the test either spins or reports a misleading "all attempts failed". Scripting one entry more than the code can consume, or using a callable that never exhausts, avoids the trap.
  • Why patch the backoff sleep instead of letting the test wait?
    Real sleeping makes the suite slow and, with jitter, non-deterministic — exactly the flakiness the test is meant to remove. Replacing the sleep function the module calls with a mock keeps the delay observable without spending it, and the recorded delay arguments let you check the backoff schedule itself, which is a genuine source of production bugs.
  • When is an iterable side_effect the wrong choice for a retry test?
    When the outcome depends on which argument was passed rather than on call position. An iterable says "the third call succeeds"; if the code batches, reorders or interleaves keys, that breaks for reasons unrelated to the behaviour under test. A callable closing over the keys it has already seen states the real rule and survives refactoring.

Rather than waiting for a dodgy turnstile to jam, you hand it a punch card that says jam, jam, open — the same rehearsal every time, over in an instant.

saying these in an interview costs you the question

  • Reproduces flakiness with real sleeps or randomness
  • Puts the success value before the exceptions
  • Assumes an exhausted iterable repeats its last item
  • Cannot explain a surprise StopIteration in a test
  • Only tests recovery, never the give-up branch
  • Raises a bare exception class the handler must inspect

context