How should a Python service's error output differ between local development, CI and production?
answer
- Three audiences, one traceback
- Colour is guessed from the destination
- Something turns the carets off
- Locals belong on a laptop, not a cluster
- Set it in the image, not the shell
basics
~20 sVary three things deliberately: colour, position anchors and captured content. Keep colour and carets for developers, force colour off where logs are stored, keep anchors everywhere unless memory is desperate, and never capture frame locals in deployment.
solid answer
~40 sTreat it as three separate switches with three separate owners. **Colour**: since Python 3.13 tracebacks are coloured when the interpreter believes it is writing to a terminal; set `PYTHON_COLORS=0` explicitly wherever output is stored, because a runner that allocates a pseudo-terminal will otherwise put escape sequences into your log index. **Position anchors**: `PYTHONNODEBUGRANGES=1` or `-X no_debug_ranges` drops the caret data and shrinks `.pyc` files and code-object memory - a bad trade in nearly every deployment, since you are paying with the diagnostic that shortens incidents. **Content**: whole-frame capture is a fine development tool and an unacceptable production default. Then make the choices structural: one entrypoint or image sets them for everyone, not eleven developers' shells, and CI asserts that a real crash reaches the log store readable.
code
console · 1 linePYTHON_COLORS=0 python3 -X no_debug_ranges -c "raise ValueError('plain traceback')"go deeper
Know that a traceback's appearance is configurable: colour can be forced off, and the caret anchors can be disabled entirely. If your output looks different from a colleague's, the environment is the first place to look.
Be able to name the switches and their effects - PYTHON_COLORS for colour since 3.13, PYTHONNODEBUGRANGES or -X no_debug_ranges for the anchors since 3.11 - and explain what each costs.
Choose by destination rather than by habit, and prove the choice: a deliberate crash whose stored log event you inspect for escape sequences, truncation and a full rendered traceback.
Frame the traceback as an interface with three consumers - developer, pipeline, future operator - and place each setting in the artefact that starts the process rather than in individual shells. Pin one interpreter version everywhere so diagnostics do not differ by environment.
An eleven-person team on a schedule-differ service will, left alone, end up with eleven slightly different diagnostic environments and one production environment nobody chose. It is worth deciding this once, because the interpreter now offers real knobs and each has a defensible setting per environment. ### The three knobs **Colour.** From **Python 3.13** the interpreter colours tracebacks - the exception type, the location line, the caret anchor - when it thinks its error stream is a terminal. `PYTHON_COLORS=0` forces it off and `PYTHON_COLORS=1` forces it on. The auto-detection is usually right, which is exactly why teams get bitten by the cases where it is not: CI runners that allocate a pseudo-terminal so that progress bars look nice, tools that re-attach a terminal to a child process, and local reproductions piped through a pager. The failure mode is not subtle but it is annoying: escape sequences embedded in stored log lines, which break substring search, confuse aggregation, and make an incident's evidence harder to read than the plain version. Decide by destination, not by inference: if the output is going to a human's terminal, colour on; if it is going anywhere that stores it, colour off, explicitly. **Position anchors.** `PYTHONNODEBUGRANGES=1`, or the equivalent `-X no_debug_ranges` at startup, tells the compiler not to keep the per-instruction column data that draws the caret line under a failing sub-expression. Tracebacks still show the source line; they stop showing which part of it failed. The saving is real - smaller `.pyc` files and less memory held by the code objects of a large codebase - and it is small. Weigh it against the thing you give up: on a line with three lookups and an arithmetic operation, the anchor is the difference between knowing the answer and reproducing the failure. In practice this switch earns its place only in genuinely memory-constrained images, and if you do set it, set it at build time as well, because a `.pyc` compiled without positions keeps its omission until it is regenerated. **Captured content.** Attaching every frame's locals to a report is a superb local debugging tool and a poor deployed default: it retains the objects those frames held and it copies whatever was in scope - tokens, connection objects, passenger records - into a store with a wider audience than the process. Development can have it. Deployment gets a chosen context dictionary and the rendered traceback. ### Where the settings should live The common failure is not choosing wrong, it is choosing in eleven places. Environment variables read by the interpreter at startup belong in the artefact that starts the interpreter: the container entrypoint, the process manager unit, the CI job definition. If a developer has to remember to export something, the setting is not policy, it is folklore. Conversely, do not push production's austerity onto developers - a local shell should have colour and anchors and may have frame capture, because the audience is one person and the process dies in a minute. ### Make it verifiable Two checks are worth automating, and both are cheap: 1. **A deliberate crash that travels the whole path.** Raise a real exception in a canary or a smoke test and assert that the stored log event contains the exception type, the file and line, and no escape sequences. This catches the pty-allocating runner, the log shipper that truncates long events, and the handler that stringified the exception and threw away the traceback. 2. **The same interpreter version everywhere.** Diagnostics have moved every release - anchors in 3.11, wider suggestions in 3.12, colour in 3.13 - so a team running one version locally and another in production will disagree about what a traceback even looks like. Pin the version in the image and use that image locally too. ### What not to do Do not build a bespoke traceback renderer to work around any of this; you will lose the anchors, the chained causes and the suggestions the interpreter computes for free. Do not turn the anchors off to "reduce log volume" - the anchor is one line and it is the useful one. Do not make the frame-capture decision per service, because the reason it is off is a property of the destination, not of the code. And do not leave the settings implicit and infer them from behaviour during an incident; a one-line startup log recording which of these are in effect costs nothing and answers the question immediately. ### The judgement being tested The interesting part is not which flag does what - it is recognising that the traceback is an interface with three different consumers: a developer at a terminal, an automated pipeline whose logs are parsed, and an operator reading an incident timeline months later. Each wants a different rendering of the same event, and the settings above are how you give it to them without writing any code.
- Colour is already disabled when output is not a terminal - why set `PYTHON_COLORS` explicitly?Because the detection is an inference about the stream, and several common setups defeat it: CI runners that allocate a pseudo-terminal for nicer output, wrappers that re-attach a terminal to a child, tools that force colour on. Escape sequences then land in stored logs, where they break search and aggregation. Explicit configuration by destination is cheaper than debugging that.
- When is `PYTHONNODEBUGRANGES=1` actually worth setting?Rarely - essentially only in memory-constrained images where the code-object and `.pyc` savings across a large dependency tree have been measured and matter. You are trading the caret anchor, which is often the fastest route from a traceback to a cause. If you do set it, set it at build time too, since a `.pyc` compiled without position data keeps that omission.
- How do you keep this policy from drifting across eleven engineers?Put the settings in the artefact that starts the interpreter - the container entrypoint, the unit file, the CI job - and use the same image locally. Anything a developer must remember to export is folklore, not policy. Then log one startup line stating which diagnostic settings are in effect, so an incident does not begin with archaeology.
saying these in an interview costs you the question
- Treats colour as harmless everywhere because terminals ignore it
- Disables caret anchors in production to save log volume
- Leaves whole-frame capture on because it helped once
- Assumes terminal detection is always right in CI
- Sets diagnostics per developer shell instead of per image
- Writes a custom traceback renderer and loses anchors and causes