How should a Python service under a supervisor signal that it is ready to take traffic?
answer
- Started is not the same as ready
- The order of start-up steps is the signal
- Listen last, or hold an explicit flag
- A backlog hang is worse than a refusal
- Readiness and liveness answer different questions
basics
~20 sDo the slow start-up work first and only then become reachable: create the listening socket last, or hold a readiness flag that the health path consults, or send the notification the supervisor is waiting for. Started is not the same as ready.
solid answer
~50 sReadiness is a promise that the process can serve a real request now, which is a different claim from `the process exists`. The cheapest signal is ordering: validate configuration, load caches and open the connections you need *before* creating the listening socket with something like `socket.create_server`, so the moment traffic can arrive is the moment you can serve it. When you must listen early — because a supervisor or a load balancer needs the port up — keep a `threading.Event` that initialisation sets, and have the readiness path answer `not ready` until it is set. Some supervisors instead pass a notification socket path in `os.environ` and wait for a datagram before declaring the unit started; the Python side of that is a few lines with the `socket` module. Whatever the mechanism, keep readiness about *this* process's initialisation, and keep it distinct from liveness.
code
python · 15 linesimport socket
def load_snapshot() -> dict[str, int]:
return {"widget-a": 4, "widget-b": 0}
def serve() -> None:
snapshot = load_snapshot() # every slow step happens before we can be reached
server = socket.create_server(("127.0.0.1", 0))
print(f"ready: {len(snapshot)} items on {server.getsockname()}", flush=True)
server.close()
serve()go deeper
Understand the distinction first: a running process is not necessarily a working one. Know that expensive start-up work belongs before the service becomes reachable, not after.
Be able to implement both mechanisms — creating the listening socket after initialisation, and a threading.Event that a readiness handler consults — and explain why a connection that hangs in the accept backlog is worse for clients than a refused connection.
Show judgement from operating real systems: keeping readiness and liveness separate, refusing to fan dependency checks into readiness, and failing start-up loudly on a bad configuration instead of sitting unready. Be ready to describe a deploy that went cold because readiness lied.
Own the convention across services: what readiness is allowed to depend on, how deploys and routing consume it, and the blast radius rule that stops a shared dependency from marking an entire fleet unready at once. Be able to defend the tradeoff against teams that want richer checks.
## The failure this prevents Consider an inventory sync service that reconciles stock between two systems, run by a small team of four. On start-up it loads a snapshot of the catalogue from the peer, builds an index and only then can answer a query. If it reports itself ready the instant the process starts, whatever routes work to it — a load balancer, a queue consumer group, a deployment that waits for `started` before stopping the old version — begins sending traffic against an empty index. Nothing crashes. The service answers, incorrectly, that it knows about no items, and a rolling deploy quietly replaces a warm process with a cold one. The bug is not in the sync logic; it is in **the claim the process made about itself**. ## Readiness versus liveness These are two different questions and conflating them is the most common mistake in this area. - **Liveness** asks `is this process wedged and in need of killing?` — its answer should almost never depend on anything outside the process. - **Readiness** asks `should work be routed here right now?` and legitimately goes false during warm-up, during a deliberate drain, or while the process is at capacity. If the same check answers both, then a slow start-up gets the process killed, or a failing dependency turns a temporary routing decision into a restart loop. ## Mechanism one: ordering The most reliable readiness signal is the one you cannot forget to send, because it is **structural**. Put every expensive step — reading and validating configuration, loading state, opening connections, compiling anything — before you create the listening socket. `socket.create_server` binds and listens in one call, so the natural shape is to call it last. Then `the port is open` and `the service works` are the same event, and no separate signal can drift out of sync with reality. A nuance worth knowing: **binding early and initialising afterwards is worse than it looks**. Once the socket is listening, the kernel completes connections into the accept backlog even though your code has not called `accept` yet. Clients do not get a clean connection refused that a retry policy handles well; they get a connection that opens and then hangs until your initialisation finishes or their timeout fires. Refusal is honest, and a hang is not. ## Mechanism two: an explicit flag Sometimes you cannot listen last — a supervisor may hand you an already-bound socket, or an orchestrator may require the health port up immediately. Then keep the state explicitly. A module-level `threading.Event` set at the end of initialisation, consulted by the readiness handler, is enough: the handler answers a failure status while the event is clear and a success status once it is set. Two rules keep it honest. - The flag must be set by the code that actually **finished** initialising, not by the code that started it. - And the readiness path must not do heavy work itself, or it becomes a **load amplifier** at exactly the moment the process is struggling. ## Mechanism three: telling the supervisor Some supervisors implement an explicit protocol: they pass the path of a **notification socket** to the child in its environment and treat the unit as started only when the process sends a datagram announcing that it is ready. Reading the variable from `os.environ` and sending one datagram is a handful of lines of stdlib socket code. The advantage over ordering alone is that the supervisor can enforce a start-up timeout and can hold back dependent units until the announcement arrives; the cost is a signal that can be forgotten and therefore lie. ## Do not put the world in your readiness check A tempting design has readiness verify every downstream dependency. It fails badly at scale: when a shared dependency has a blip, every process in the fleet reports itself unready at once, everything is removed from rotation simultaneously, and a partial degradation becomes a total outage. Keep readiness to **what this process controls** — initialisation finished, capacity available, not draining — and handle dependency failures inside the request path with timeouts, retries and a fallback. The clock-skew case that plagues an inventory sync is a good example of the distinction: if the peer's timestamps arrive minutes ahead of local time, that is a data problem to log, measure and skip records over, not a reason for the process to declare itself unready and take itself out of rotation. ## And fail start-up loudly If a hard prerequisite is missing — an unreadable credential, a configuration key with no value — the answer is not to start and sit there unready forever. Write the reason to `sys.stderr` and **exit non-zero**, so the supervisor's restart policy and the deployment tooling both see a failure instead of a process that is up and useless.
- Why is a process being up not the same as the service being ready?Because the interpreter starting says nothing about whether configuration validated, caches loaded or connections opened. A process can be running and answer every request wrongly or slowly. Readiness is a claim the application makes about its own initialisation; process existence is a claim the operating system makes, and only the first one should gate traffic.
- What happens to clients if the socket is listening but initialisation has not finished?The kernel completes their connections into the accept backlog, so from the client's side the connection succeeds and then hangs with no response until initialisation finishes or the client times out. That is worse than a connection refused, which a retry policy handles immediately and which makes the state of the process obvious.
- Should the readiness path check downstream dependencies?Rarely. If every process checks a shared dependency, one blip makes the whole fleet report unready at the same moment and removes everything from rotation, turning partial degradation into an outage. Keep readiness scoped to this process — initialisation done, capacity available, not draining — and deal with dependency failures in the request path with timeouts and fallbacks.
Unlocking the shop door is the readiness signal customers actually respond to; propping it open while you are still unpacking the shelves means people walk in and leave empty-handed.
saying these in an interview costs you the question
- Treats a started process as a ready process
- Binds the listening socket before loading state
- Uses one endpoint for both liveness and readiness
- Reports ready and serves errors until warm
- Checks every downstream dependency in the readiness path
- Starts unready forever instead of exiting on bad configuration