skip to content

Why is a doctest brittle when its example prints a float, a set or a default repr?

level: middleimportance: should knowfreq 34%

answer

  1. The comparison is text, not value
  2. Representation error is visible in the repr
  3. Hash order changes between processes
  4. Print something canonical before adding flags
  5. No numeric tolerance flag exists

basics

~20 s

Because doctest matches printed text exactly. 0.10 + 0.20 prints 0.30000000000000004, set order depends on the hash seed, and a default repr embeds a memory address, so the example fails for reasons unrelated to the code being correct.

solid answer

~40 s

The assertion is a string comparison, so anything whose printed form is not deterministic produces false failures. Binary floating point makes `0.10 + 0.20` print `0.30000000000000004` in an ad-auction bidder; a `set` repr is in hash order and varies with `PYTHONHASHSEED` between runs; a default `repr` carries the object's address. The first fix is to print something canonical — `round(price, 2)`, an f-string format, `sorted(...)` — rather than to reach for a flag. Where the variation is genuinely irrelevant, `ELLIPSIS` lets `...` match any text, `NORMALIZE_WHITESPACE` collapses whitespace differences from wrapped output, `IGNORE_EXCEPTION_DETAIL` tolerates changing exception messages, and `# doctest: +SKIP` disables an example entirely. There is no numeric-tolerance flag. If an example needs several flags to pass, it is printing the wrong thing and the real assertion belongs in a dedicated test.

code

python · 16 lines
python
import doctest


def bid_increment(bid, step):
    """Raise a bid by one step.

    >>> bid_increment(0.10, 0.20)
    0.30000000000000004
    >>> round(bid_increment(0.10, 0.20), 2)
    0.3
    """
    return bid + step


if __name__ == "__main__":
    print(doctest.testmod())

go deeper

for a junior

Recall that the check is on printed text, so 0.1 + 0.2 in an example must be written exactly as the interpreter prints it. Knowing to round or format the value in the example is enough at this level.

for a middle

Explain each source of instability — representation error, hash-ordered sets, address-bearing reprs, changing exception messages — and name the flag or the rewrite that addresses each. Know that dicts are ordered and sets are not.

for a senior

Demonstrate the judgement: treat flags as a smell to be justified, insist that an example prints something a reader wants to see, and recognise when the assertion actually needs tolerance and therefore belongs outside the docstring.

for a principal

Own the house rule that stops a suite decaying into ellipses: which flags are allowed run-wide, which need a reason in review, and how you keep verified examples valuable as documentation rather than a second, weaker test suite.

### The mechanism: a string comparison in a language with unstable strings doctest compares the text an example printed to the text you wrote beneath it, character for character. So every source of *non-determinism in printed form* is a source of false failures. Three families cover almost all real cases, and a fourth is worth knowing. #### 1. Floats: representation error is visible in the repr Binary floating point cannot represent most decimal fractions exactly. In an ad-auction bidder, raising a bid of `0.10` by a step of `0.20` does not print `0.3`: ```pycon >>> 0.10 + 0.20 0.30000000000000004 ``` Since 3.1 CPython prints the *shortest* string that round-trips to the same double, which makes the repr stable and deterministic — but it also makes the accumulated drift visible. An example written as `0.3` fails, and worse, the same computation reached by a different order of operations may print a *different* wrong-looking value, so the docstring becomes coupled to an implementation detail of the arithmetic. The fixes, in order of preference: print a rounded value in the example (`round(price, 2)`), format it explicitly (`f"{price:.2f}"`), or move the assertion out of the docstring into a test that uses `math.isclose` with an explicit tolerance. doctest has **no** numeric-tolerance flag; anyone who claims otherwise is thinking of a third-party runner's extension. #### 2. Hash-ordered containers A `set` or `frozenset` repr lists its elements in hash order, and for `str` keys that order depends on the per-process hash seed (`PYTHONHASHSEED`), so it changes between runs: ```pycon >>> {"floor", "cap", "bid"} {'cap', 'floor', 'bid'} ``` Run it again and the order may differ. The fix is to print something canonical — `sorted(...)` — not to chase a flag. Note the asymmetry: **dicts are safe**. Insertion order has been a language guarantee since 3.7, so a dict literal's repr is deterministic; a set's is not. #### 3. Reprs that embed identity or environment The default `object.__repr__` contains the object's address, and many built-in reprs embed one (`<generator object bids at 0x7f...>`). Paths, timestamps, host names and temporary directories are the same class of problem. Here `ELLIPSIS` genuinely earns its place: with the flag on, the marker `...` in expected output matches any run of text. #### 4. Tracebacks doctest already ignores the *body* of an expected traceback and compares only the header line and the final `ExceptionType: message` line — so caret and anchor lines (added in 3.11) and frame details cost you nothing. What still bites is the **message**: exception messages are not API and do change between versions and between call paths. `IGNORE_EXCEPTION_DETAIL` reduces the comparison to the exception class. ### The flags, and what each is honestly for | flag | what it does | when it is right | |---|---|---| | `ELLIPSIS` | `...` in expected output matches any text | addresses, paths, long or variable output | | `NORMALIZE_WHITESPACE` | all runs of whitespace compare equal | output you wrapped for readability | | `SKIP` | the example is parsed but never executed | purely illustrative lines, or a known-slow call | | `IGNORE_EXCEPTION_DETAIL` | compares only the exception class | messages that vary by version | | `REPORT_NDIFF` | shows a character-level diff on failure | debugging a near-miss you cannot see | Set them per example with a directive comment (`# doctest: +ELLIPSIS`), or run-wide with the `optionflags` argument or the `-o` CLI option. ### The judgement an interviewer is actually listening for Flags are anaesthetic, not treatment. `ELLIPSIS` on a whole line asserts almost nothing — an example reading `>>> price(bid)` / `...` has stopped being a test. The rule of thumb worth saying out loud: **if the example needs more than one flag to pass, the value being printed is the wrong value to print.** Rewrite the example to produce something canonical and readable — a rounded number, a sorted list, an explicit format string — and if the real assertion needs tolerance, ordering tolerance, or a fixture, it was never a documentation example in the first place and belongs in a dedicated test.

  • When is ELLIPSIS the right fix, and when is it hiding the assertion?
    It is right when the varying text is genuinely irrelevant — an object address, a temporary path, a long tail of output whose shape you have already checked. It is hiding the assertion when `...` covers the value under test: an example whose expected price is an ellipsis passes for every price, so it has stopped testing anything and has stopped documenting anything too.
  • Why is a dict literal safe in a doctest today while a set is not?
    Dict iteration and repr follow insertion order, a language guarantee since 3.7, so a dict's printed form is deterministic. A set is stored by hash, and for str elements the hash is randomised per process by default, so its repr order can change between runs. Print `sorted(the_set)` when you need it in an example.
  • What exactly does the `# doctest: +SKIP` directive do to an example?
    The example is still parsed and still displayed to a reader, but it is never executed and never compared — the run counts it as skipped. It is honest for a line that is illustrative, needs credentials, or is slow, and dishonest as a way to silence a failure that reflects a real defect.

Pinning a doctest to a raw float or a set repr is like proofreading a photocopy against the original: any smudge the machine adds counts as a difference, however faithful the text is.

saying these in an interview costs you the question

  • Thinks doctest compares floats with a tolerance
  • Claims Python's float addition is simply broken
  • Adds ELLIPSIS everywhere instead of printing stable values
  • Assumes set repr order is stable between runs
  • Expects doctest to compare tracebacks frame by frame
  • Thinks NORMALIZE_WHITESPACE also ignores case and punctuation

context