In a web framework's middleware chain, why is the correlation-id hook placed first, ahead of logging and authentication?
answer
- nothing downstream can invent it
- first in, before any record
- adopt inbound or generate fresh
- per-request context, not a global
- returned on the response, sent onward
basics
~20 sA correlation id must exist before anything downstream writes a log line, records a metric or calls another service, so the hook that adopts or generates it runs first. Anything registered ahead of it emits records nothing can join.
solid answer
~50 sThe correlation id is the one piece of per-request state every other hook depends on, so it goes at the very outside of the chain. The hook does three things: read an inbound request-id header if the caller sent one and it passes a strict shape check, otherwise generate a fresh value; publish it into the per-request context that the logging layer and outbound clients read; and return it on the response so a caller can quote it. Put it first and every later record carries it — the access log line, the throttled rejection, the authentication failure, the stack trace from the error stage. Put it after authentication and you lose exactly the requests you most want to trace: the `401`s, the `429`s and the malformed bodies that never reach a handler. It is cheap, it does no I/O, and it never rejects a request.
go deeper
Remember the shape: one id per request, established by the outermost hook, attached to every log line and echoed on the response. Be able to say why a handler is too late a place to create it.
Explain the adopt-or-generate decision, why the value goes into per-request context rather than a parameter or a global, and which categories of request lose their id when the hook is placed after authentication.
Show you have debugged the gaps: ids missing from asynchronous work, unvalidated caller values landing in every log line, and an edge layer that mints an id the application then overwrites instead of adopting.
Frame it as a platform contract: which layer mints the id, which layers must adopt it unchanged, what the shape rule is, and how you keep that contract from drifting across services written by different teams.
## What the correlation-id hook is A **correlation id** is one opaque value that identifies a single inbound request for that request's whole lifetime. Every log line written while handling it, every metric sample it contributes to, and every call the service makes to another service carries the same value, so an operator holding one id can pull back everything that happened. The element that establishes it is an ordinary member of the framework's middleware chain: it runs on the way in, delegates to the rest of the chain, and does a little work on the way out. Concretely it does four things: 1. **Adopt or generate.** If the request carries a request-id header and the value passes a strict shape check, reuse it — that is what joins this service's records to the caller's. Otherwise generate a fresh value. 2. **Publish it.** Write the value into whatever per-request context the framework offers, so that a log call ten frames deep picks it up without anyone passing it as an argument. 3. **Propagate it.** Outbound calls made while handling the request send the same id on, so the next service adopts rather than invents one. 4. **Return it.** Put the id on the response, so a user quoting an error page or a client logging a failure hands you the exact search key. ## Why first, and not merely early The rule is mechanical: **a consumer cannot read state its producer has not yet written.** Every element of the chain that emits a record is a consumer — the access log, the metrics hook, the rate limiter when it rejects, the authentication hook when it fails, the error stage when it catches. If any of them is wrapped *outside* the correlation hook, its records are written before the id exists and are unjoinable forever. The requests where placement shows are the ones that never reach a handler: - a request refused by a body-size or content-type guard - a request answered `401` because its credential had expired - a request answered `429` by a limiter - a request that failed inside an inner hook and unwound to the error stage These are usually the most interesting requests during an incident, and they are precisely the ones a late-placed id hook drops. | Correlation hook placed | Records that carry the id | Records that do not | |---|---|---| | Outermost | access log, metrics, limiter rejections, auth failures, handler, error stage | none inside the chain | | After authentication | handler, error stage, anything the handler calls | access log for rejected requests, limiter rejections, auth failures | | Inside the handler | log lines the handler writes | everything the framework itself emits | ## Adopting an id from an untrusted caller An inbound id is caller-controlled input, and it flows into every log line and every outbound header — so it needs a cheap guard: a length cap and a restricted character set, with a generated id as the fallback. Two further points come up in interviews: - **An id is a join key, never an authorization token.** It must not grant access, address a resource, or be treated as unguessable. - **If you reject the caller's value, say so.** Recording the rejected value in a separate field keeps the join available to a human without letting unvalidated text become the primary id. ## Where the id has to live The id belongs in **per-request context** — state the framework scopes to one in-flight request — not in a process-wide variable and not on a long-lived shared object. A server handles many requests at once; anything shared between them will interleave, and the symptom is the worst kind: log lines attributed to the wrong request. Frameworks differ here, and it is worth saying so honestly in an interview: some carry per-request context across asynchronous boundaries automatically, so work handed to another pool keeps the id; others require the code doing the hand-off to capture the context explicitly, and forgetting is why an id present in the access log is missing from a background task's logs. ## What good placement looks like - Outermost of the application's own hooks — nothing meaningful sits outside it. - Cheap: no I/O, no lookups, no allocation worth worrying about. - Never short-circuits: it has no reason to reject anything. - Adopts rather than regenerates when an edge layer already minted an id. - Returns the id on **every** response, including failures. Get this one placement right and every other observability decision in the chain has something to hang on. Get it wrong and each of them degrades quietly, in a way that only shows up during the incident when you can least afford it.
- What should the hook do with a correlation id supplied by an untrusted caller?Validate it cheaply — a length cap and a restricted character set — and adopt it only if it passes; otherwise generate a fresh id and, if the join matters, record the caller's value in a separate field. The id is a join key, so it must never be treated as a secret or used to authorize anything.
- The id appears in the access log but is missing from log lines written by a background task the handler started. Why?The id lives in per-request context, which is bound to the execution that the framework started for the request. Handing work to another pool can leave that context behind. Some frameworks propagate it across asynchronous boundaries; where they do not, the code doing the hand-off must capture the id explicitly and re-establish it.
- Does it matter whether the correlation hook or the access-log hook is outermost?Yes. The access log is the outermost observer, but it still sits inside the correlation hook, so the line it writes on the way out already carries the id. Reverse the two and every access-log line — including those for rejected requests — is written without one.
It is the case number a clerk writes on the folder before anyone files a single page into it. Pages filed before the number exists can never be matched back to the case.
saying these in an interview costs you the question
- Generates the id inside the handler, after outer hooks have already logged
- Assumes every framework attaches a request id automatically
- Reflects a caller-supplied id with no length or character-set check
- Places the hook after authentication, losing all rejected requests
- Keeps the id in a process-wide variable shared by concurrent requests