skip to content

How do you write a `__repr__` that stays useful in production logs and debuggers?

level: seniorimportance: should knowfreq 38%

answer

  1. Read at the worst possible moment
  2. Cheap, total, honest
  3. Never raise, never leak, never block
  4. Truncation must announce itself
  5. reprlib for bounded and recursive output

basics

~20 s

Show the class name and the few fields that identify the instance, keep it cheap, side-effect free and unable to raise, redact secrets, and bound large fields while printing the real size next to any truncated sample.

solid answer

~50 s

A `__repr__` is read at the worst possible moment — a failing assertion, a log line, a debugger paused in a live process — so it has to be cheap, total and honest. Cheap: no I/O, no lazy loading, no computation you would not run in a tight loop, because a debugger re-evaluates the repr of every visible object on each step. Total: it must not raise, so read risky attributes with a `getattr` default rather than assuming a fully-built instance. Honest: if you bound a large field with `reprlib.repr`, print the real length beside the sample, or a truncated list reads as a short one and people debug the wrong thing. Redact credentials explicitly rather than hoping they are never logged. Prefer `Class(field=value, ...)` for reconstructible objects and `<Class state>` for anything holding live resources.

code

python · 21 lines
python
import reprlib


class ExtractedVideo:
    def __init__(self, path, frames, api_token):
        self.path = path
        self.frames = frames
        self._api_token = api_token

    def __repr__(self):
        return (
            f"{type(self).__name__}(path={self.path!r}, "
            f"frame_count={len(self.frames)}, "
            f"sample={reprlib.repr(self.frames)}, token=<redacted>)"
        )


v = ExtractedVideo("clip.mp4", list(range(4800)), "sk-secret")
print(repr(v))
# ExtractedVideo(path='clip.mp4', frame_count=4800,
#                sample=[0, 1, 2, 3, 4, 5, ...], token=<redacted>)

go deeper

for a junior

Start with the habit rather than the policy: give each class a one-line __repr__ showing the class name and the constructor fields, and keep anything slow or secret out of it. That alone makes your objects readable everywhere.

for a middle

Explain the constraints and the tools: no side effects, no exceptions, bounded output with reprlib.repr, cycle protection with reprlib.recursive_repr(), and type(self).__name__ so subclasses report themselves correctly.

for a senior

Demonstrate incident thinking. Talk about a repr being evaluated by a debugger on every step, about truncation that hides the real size, about redaction at the source, and about designing the text so it can be grepped and quoted in a runbook.

for a principal

Own the convention across teams: an agreed repr shape, a redaction marker, a rule that reprs are cheap and total, and the recognition that log and incident quality depends on reprs being consistent and stable rather than individually clever.

Take a concrete setting: a video-metadata extractor maintained by a four-person team, running as a long-lived worker. Every object it handles ends up in a log line or in a debugger session eventually, and the only thing standing between an engineer and the truth about that object is its `__repr__`. Designing that method is a production concern, not a formatting nicety. ## Rule one: it must be honest about what it hides The failure that teaches this is a silent truncation. Someone writes a repr that shows the first six decoded frames because the full list is enormous, and it reads `Extracted(frames=[0, 1, 2, 3, 4, 5, ...])`. Months later a batch produces short clips, an engineer sees a similar-looking line, and concludes the extractor is dropping frames — the truncation and the real defect are indistinguishable in the log. - The fix is one field: print the true count next to the sample, `frame_count=4800, sample=[0, 1, 2, 3, 4, 5, ...]`. - `reprlib.repr` gives you the bounded rendering with sensible limits, and `reprlib.Repr` lets you configure them, but the count is yours to add. A repr that abbreviates must say that it abbreviated and by how much. ## Rule two: it must be cheap A debugger evaluates the repr of every object in scope, on every step. A watch expression does it repeatedly. Logging does it once per record, sometimes at a high rate. So a `__repr__` that opens a file, queries a database, triggers a lazy load or computes a hash over a large buffer converts debugging from a diagnostic activity into a source of load — and worse, changes the behaviour of the program you are inspecting. If a value is expensive, show a cheap proxy: a length, an id, a state flag. ## Rule three: it must not raise Reprs are called on objects in states you never designed for: - a half-constructed instance whose `__init__` blew up mid-way, - an object being reported inside an exception handler, - an object resurrected from a pickle with fields missing. If `__repr__` raises, the error message you needed is replaced by a second, unrelated traceback, and the original context is lost. Read anything uncertain with `getattr(self, 'field', '<unset>')`. For self-referential structures, decorate with `reprlib.recursive_repr()`, which substitutes `...` when the same object is re-entered — the same trick the built-in containers use on themselves. ## Rule four: it must not leak Whatever the repr contains goes into logs, crash reports, and any error-tracking pipeline you have. Tokens, passwords, connection strings with embedded credentials and personal data all belong behind an explicit `<redacted>` placeholder. Putting the redaction in `__repr__` is stronger than a logging filter, because it protects every path that renders the object, including ones added later by someone who never read the filter config. ## Rule five: it should carry identity plus a few stable fields - The useful content is the class name — use `type(self).__name__` so subclasses do not lie about themselves — then the fields that let you tell this instance apart from its neighbours: the input path, an id, a status, a size. - Render each with `!r` so strings are quoted. - Keep the whole thing on one line; a repr that wraps is unreadable in a grepped log. Anything larger belongs in an explicit diagnostic method the caller opts into, not in the repr that fires automatically. ## Rule six: treat it as a shared convention With four people on a service, the value of reprs comes from their consistency: same shape, same field ordering, same redaction marker, same abbreviation style. That is worth a short written rule and a review habit, because the payoff is collective — one person's incident is read using everyone else's reprs. - Choose the reconstructible `Class(field=value)` form for value objects, - and the `<Class state>` angle-bracket form for objects holding live resources, so the shape itself tells a reader whether the text could be evaluated. Finally, keep the repr stable over time: people learn to scan for a familiar shape, and changing the field order or names quietly invalidates saved log queries and every incident runbook that quotes one.

  • How do you keep `__repr__` from recursing forever on a self-referential object?
    Decorate it with `reprlib.recursive_repr()`, which detects re-entry for the same object on the same thread and substitutes `...` instead of recursing. The built-in containers do the equivalent internally, which is why a list appended to itself prints `[[...]]` rather than blowing the stack.
  • What belongs in `__repr__` versus a dedicated diagnostic method?
    The repr gets identity and a handful of cheap, stable fields on one line, because it fires automatically wherever the object is rendered. A full dump — every field, decoded payloads, computed statistics — belongs in an explicit method the caller opts into, where cost and verbosity are the caller's choice rather than a surprise in a log.
  • Why put credential redaction in `__repr__` rather than in a logging filter?
    Because the repr protects every rendering path, not just the one you configured: exception messages, debugger panes, a printed list of objects, an error-tracking payload. A filter covers the paths you remembered; the repr covers the ones a future colleague adds. Keeping the secret out of the text at the source is the stronger boundary.

saying these in an interview costs you the question

  • Doing I/O or a lazy load inside __repr__
  • Dumping an unbounded collection into the repr
  • Leaking tokens or passwords into log output
  • Letting __repr__ raise on a half-built object
  • Truncating a field without showing its real size
  • Treating __repr__ as user-facing product copy

context