skip to content

How do you write a logging.Filter that redacts a token from every LogRecord?

level: middleimportance: should knowfreq 42%

answer

  1. One method decides the record's fate
  2. The return value keeps, drops or replaces
  3. Rewrite the parts before a formatter runs
  4. Clear args after collapsing the message
  5. Handler.addFilter, not the parent logger

basics

~20 s

Give an object a filter(record) method, or pass any callable taking a LogRecord, rewrite record.msg and record.args in place, and return True to keep the record. Attach it with Handler.addFilter so every record reaching that destination is scrubbed.

solid answer

~40 s

A filter is any object with a `filter(record)` method; subclassing `logging.Filter` is only a convenience, and since 3.2 a bare callable works too. Returning a falsy value drops the record entirely, `True` keeps it, and since 3.12 you may return a `LogRecord` to substitute a scrubbed copy rather than mutating the original. To redact, rewrite `record.msg` and `record.args` before any formatter runs; if you collapse them with `record.msg = record.getMessage()`, clear `record.args` too, or the next `getMessage()` will try to interpolate the finished string. Placement is what people get wrong: `Logger.addFilter` only sees records logged *directly* on that logger, so a filter on `"svc"` never sees a record from `"svc.db"`, which merely propagates to the ancestor's handlers. Redaction belongs on the handler, where every record must pass.

code

python · 19 lines
python
import logging
import re

SECRET = re.compile(r"(?i)\b(api_key|token|password)=\S+")


class Scrub(logging.Filter):
    def filter(self, record: logging.LogRecord) -> bool:
        record.msg = SECRET.sub(r"\1=<redacted>", record.getMessage())
        record.args = ()
        return True


handler = logging.StreamHandler()
handler.addFilter(Scrub())
log = logging.getLogger("app")
log.addHandler(handler)
log.setLevel(logging.INFO)
log.info("login failed token=%s", "tok_abc123")  # login failed token=<redacted>

go deeper

for a junior

Know that logging has a filter hook at all, and recall its shape: an object with a filter(record) method that returns True to keep the record. Be able to attach one with addFilter without looking it up.

for a middle

Explain the mechanics: which record attributes hold the template and the values, why args must be cleared after you collapse them, and what a falsy return does. Be ready to write a small filter on the spot.

for a senior

Demonstrate the placement judgment — handler, not ancestor logger — and name what redaction still misses: extra= attributes, values hidden inside an object's repr, and formats the regex was not written for. Treat filter hits as a signal that a call site needs fixing.

for a principal

Own it as a control with known limits. Decide whether redaction is a backstop or the primary defence, whether it fails loudly or open, how it is declared once in shared configuration rather than per service, and how you measure the call sites it keeps catching.

## The protocol A logging filter is not a class you must inherit from. Anything with a `filter(record)` method qualifies, and since Python 3.2 a plain callable taking a `LogRecord` does too, so `lambda record: "tok_" not in record.getMessage()` is a legal filter. `logging.Filter` itself exists mostly for its one built-in behaviour: constructed with a name, it passes only records from that logger or its descendants. For redaction you subclass it (or write a callable) and ignore that behaviour. The return value has three meanings. Falsy means *drop this record* — no handler downstream of that point will see it. `True` means keep it. Since 3.12, returning a `LogRecord` instance means *use this record instead*, which lets you build a scrubbed copy rather than mutating the caller's data — useful because the same record object is shared by every handler on the propagation path. ## What to rewrite The interesting attributes are `record.msg` (the format string as passed) and `record.args` (the values). If the call site used lazy `%`-style arguments, the cleanest redaction is structural: walk `record.args` and replace the elements that are credentials. If the call site interpolated eagerly you are left with prose, and the only option is a regex over `record.getMessage()`. A common shape collapses both: ```python record.msg = SECRET_RE.sub("<redacted>", record.getMessage()) record.args = () ``` Clearing `args` is not optional. `getMessage()` runs `msg % self.args` whenever `args` is truthy, so leaving the old tuple in place means the finished text is interpolated a second time — which raises on a stray `%` or silently corrupts the line. Do not stop at the message. A `LogRecord` also carries whatever the call passed as `extra=`, set directly as attributes on the record, and a formatter whose format string names them will print them untouched. `record.name`, `record.pathname`, `record.funcName` and the exception text attached to the record are all rendered by some format strings. If your redaction only looks at the message, an `extra={"authorization": token}` walks straight past it. ## Where to attach it — the mistake that matters `Logger.addFilter` and `Handler.addFilter` sound symmetric and are not. When you log on `"svc.db"`, `Logger.handle` applies *that logger's* filters and then calls `callHandlers`, which walks up the ancestor chain invoking the **handlers** it finds. Ancestor loggers' filters are never consulted. So a redaction filter attached to `"svc"`, or even to the root logger, silently misses every record produced by a child logger — which, in a package that follows the `getLogger(__name__)` convention, is essentially all of them. A filter on a handler sees every record that reaches that handler, whichever logger produced it. That is the placement for a security control. The cost is that you must attach it to each handler, which is exactly what `logging.config.dictConfig` is for: declare the filter once in the `filters` section and list its key under each handler's `filters`. ## Failure modes to expect A filter that raises does not fail quietly. Filters run outside the `try` that routes emit errors to `Handler.handleError`, so an exception inside `filter()` propagates out of the `log.info(...)` call and into the application. For a redaction filter that is arguably the safer default — a broken scrubber that returns nothing would leak — but it must be a deliberate choice, and the filter body should be simple enough to trust. A filter is also a blunt instrument for a fundamentally lossy job. Regexes over free text miss formats they were not written for, and mangle innocent text that resembles one; a token that reaches the record inside a nested dict's `repr` is not going to be matched at all. Redaction at the log boundary is a backstop, not a substitute for not passing secrets to log calls in the first place. Treat a hit in the filter as a signal that some call site needs fixing, and consider having the filter count its hits so that signal is visible. ## Related knobs `logging.setLogRecordFactory` replaces the callable that builds every record, which is the place for adding attributes rather than removing them. `logging.LoggerAdapter` lets a caller inject contextual data on the way in. Neither is a redaction point: they run at record creation, before you know which handler will render what.

  • Why attach the redaction filter to the handler rather than to the top-level logger?
    Because ancestor loggers' filters are never applied to a child's record. Logger.handle runs only the filters of the logger the call was made on, then callHandlers walks the ancestors invoking their handlers. With the usual getLogger(__name__) convention almost every record comes from a child, so a filter on the parent logger scrubs nothing. A handler filter sees every record that reaches that destination.
  • Your filter rewrites record.msg but a token still appears in the output. What did you miss?
    Most likely fields outside the message. Values passed as extra= become plain attributes on the record and are printed verbatim by any format string that names them. Leaving a stale record.args in place can also reinterpolate the text. And if the value only appears inside an object's repr, a regex over the message will never match it.
  • What happens if a filter raises an exception?
    It propagates into the caller of the logging call. Filters run outside the try/except that sends emit failures to Handler.handleError, so a bug in a filter breaks the code that logged. For a redaction filter, failing loudly beats failing open, but keep the body simple and cheap: it runs on every record that reaches the handler.

saying these in an interview costs you the question

  • A filter on a parent logger also scrubs child loggers' records
  • Returning False from filter() just skips the redaction step
  • Filters run after the formatter has built the output line
  • Rewriting record.msg alone always removes the secret
  • Only a logging.Filter subclass can be used as a filter
  • Filters cannot drop a record, only levels can

context