Why is logging.config.dictConfig input a trust boundary, and would you ever run logging.config.listen in production?
answer
- Ask what the configurator does with a section
- Some keys name things to import and call
- The older INI reader evaluates strings
- A socket that applies what it is sent
- Translate an in-house schema, never pass one
basics
~20 sBoth apply configuration as code. dictConfig's "()" and "class" keys name dotted paths that it imports and calls at configure time, and listen() applies whatever arrives on a socket. Treat logging configuration as reviewed source, never as data from outside.
solid answer
~50 s`logging.config.dictConfig` is not a data reader. A `"()"` key names a callable that it resolves by dotted path and *calls*, with the section's remaining keys as keyword arguments; a handler's `"class"` key is imported and instantiated the same way. Whoever controls that dictionary controls code execution in the process, at process privilege. `logging.config.fileConfig` is worse still: it evaluates its `args` and `kwargs` strings with `eval()`. `logging.config.listen()` layers a socket on top — it starts a thread serving length-prefixed payloads and applies each one, JSON through `dictConfig` and anything else through `fileConfig`. It binds loopback and takes a `verify` callable (since 3.4) that can reject bytes before they are used, but the honest answer for production is no: ship the configuration with the code, and if it must change at runtime, re-read a vetted file on a signal rather than accepting one over a socket.
code
python · 18 linesimport logging
import logging.config
def build_filter():
print("factory executed while the configuration was applied")
return logging.Filter()
logging.config.dictConfig(
{
"version": 1,
"filters": {"scrub": {"()": "__main__.build_filter"}},
"handlers": {"out": {"class": "logging.StreamHandler", "filters": ["scrub"]}},
"root": {"handlers": ["out"], "level": "INFO"},
}
)
logging.getLogger("job").info("configured")go deeper
Know that logging is configured from a dictionary via logging.config.dictConfig, and that this configuration belongs in the repository next to the code rather than somewhere it can be edited at runtime.
Explain what the configurator does with a section: a "()" or "class" key is a dotted path it imports and calls, so the dictionary can construct arbitrary objects. Be able to say why that makes its source matter.
Show the threat model concretely — a writable config file, an environment-selected path, a fetched fragment — and the operational cousin where a stale configuration leaves a service at DEBUG dumping payloads. Know what verify does and why loopback is not a boundary.
Own the policy: configuration ships with the code, runtime tuning goes through a narrow in-house schema your own code translates into dictConfig, and the socket listener is declined outright. Be ready to defend that against a team that wants live reconfiguration, and to make the applied configuration observable and audited.
## The configuration is a program It is natural to think of a logging configuration as inert settings — levels, formats, file paths. It is not. `logging.config.dictConfig` supports a user-defined object syntax: a section containing the key `"()"` names a factory, given either as a callable or as a dotted string that the configurator resolves by importing it, and the remaining keys of that section become keyword arguments to a call. Handlers use `"class"` the same way. Filters, formatters and handlers can all be constructed this way, and the call happens the moment `dictConfig` runs. So the answer to "what can someone do who controls this dictionary?" is: import any module reachable on the interpreter's path and call it with arguments of their choosing, inside your process, with your privileges, usually before the application has finished starting. That is not a logging bug — it is the documented extension mechanism, and it is why extensible logging exists at all. It just means the dictionary is code, and inherits the trust rules of code. `logging.config.fileConfig` is the older INI-style entry point and is sharper: it passes handler `args` and `kwargs` strings through `eval()`. There is no interpretation of that which is safe for untrusted input. ## The socket channel `logging.config.listen(port=9030, verify=None)` returns a thread that, once started, serves a tiny protocol: a four-byte big-endian length followed by that many bytes of configuration. The receiver tries `json.loads`; a dict goes to `dictConfig`, anything else is treated as a file and goes to `fileConfig`. `stopListening()` shuts it down. Two mitigations exist. The server binds `localhost`, so it is not reachable from another host by default — though "local" is a weak boundary in a container with a port-forward, on a shared build agent, or anywhere another process runs as a different user on the same box. And `verify`, added in 3.4, is a callable given the raw received bytes that returns `None` to reject them or the bytes (possibly transformed, for instance decrypted) to use. With an HMAC check in `verify` the channel is defensible. Defensible is not the same as worth having. The feature buys runtime reconfiguration of one process; the same need is met by re-reading a file you already control on `SIGHUP`, or by an admin endpoint that goes through your existing authentication and only ever adjusts levels. Those have a bounded blast radius; `listen` has the blast radius of `exec`. On any service I own, that trade does not clear. ## Where the untrusted dictionary actually comes from The interesting exposure is rarely someone shouting at port 9030. It is the ordinary supply path: - a configuration file in a directory the application user can write, so any file-write bug becomes code execution at the next restart; - a path chosen by an environment variable, in a deployment where environment is easier to influence than code; - a configuration fetched from a central service and merged into the dictionary, where the security of your process now equals the security of that service and the channel to it; - a fragment supplied per-tenant or per-team and merged in, which is untrusted input wearing a settings hat. There is a second, quieter failure in that last shape. A stale cached copy of a fetched configuration — served after the source has moved on — can silently leave a service at `DEBUG` with a handler that dumps request payloads. At a 1,200-request-per-minute peak that is a disclosure incident measured in minutes, and nothing about it looks like an error: the process is doing exactly what its configuration says. ## What I would actually require Ship the configuration with the code. It lives in the repository, it is reviewed like code, and it is baked into the artefact, so the deployed dictionary has the same provenance as the deployed program. Where something must be tunable at runtime, do not pass a foreign dictionary through. Define a small in-house schema — level per logger, sample rate, destination chosen from a fixed set — validate it, and *translate* it into the `dictConfig` structure yourself. The translation layer is the trust boundary: `"()"` and `"class"` keys never come from outside, because outside cannot express them. Also remember that logging configuration is process-global and that `dictConfig` replaces the world by default; a merge that half-applies leaves a service with no destination at all, which is its own outage. Make the applied configuration observable — log its source and a digest of it at startup — and give reconfiguration an audit trail, because a change in where the logs go is exactly the change an attacker wants and exactly the change nobody reviews.
- If a team insists on runtime reconfiguration, what would you accept instead of logging.config.listen?A vetted file the service re-reads on a signal, or an admin endpoint behind the service's existing authentication that accepts only a small validated schema — level per logger, sample rate, a destination chosen from a fixed set — which the service then translates into a dictConfig structure itself. The blast radius becomes a wrong log level rather than arbitrary code, and the change is authenticated and auditable.
- How do you let teams contribute logging configuration without handing them code execution?Do not accept their dictionary. Publish a narrow schema they can fill in, validate it against an allowlist, and generate the dictConfig structure in shared code where the "()" and "class" keys are written by you. The translation layer is the trust boundary; if outside input cannot express a factory or a class path, a hostile fragment has nothing to reach for.
- What is the operational risk of a stale logging configuration, separate from code execution?It can leave a service at DEBUG with a handler that records request payloads long after someone believed the change was reverted, which is a disclosure incident that looks like normal operation. It can also point handlers at a destination nobody is watching, so an outage is invisible. Log the configuration's source and a digest at startup, and treat a change in where logs go as an audited change.
saying these in an interview costs you the question
- A logging configuration is just data, so it is safe to load
- Only the socket listener is risky; a config file is inert
- localhost binding makes logging.config.listen safe enough
- verify only checks that the configuration parses
- Merging a fragment from another team carries no privilege
- The worst case is logs going to the wrong file