skip to content

How do you use a correlation or request id in HTTP API error responses so that support can trace a reported failure, and where should that id come from?

level: middleimportance: should knowfreq 48%

answer

  1. assign at the edge, generate if absent
  2. request-scoped context so all logs carry it
  3. header on all responses, body on errors
  4. propagate downstream: one key, whole graph
  5. validate inbound id: log injection, cardinality

basics

~20 s

Generate or accept one id per inbound request, put it in every log line and in every error body (and ideally a response header), and propagate it to downstream calls. Support then searches logs by the id the caller quotes.

solid answer

~50 s

Assign an id at the edge: accept a caller-supplied one if present and valid, otherwise generate it. Store it in a request-scoped context so every log line for that request carries it automatically, echo it on every response — a header for all responses and inside the body for errors — and forward it on outbound calls so downstream services log the same id. That gives a single search key spanning the whole call graph, which is what makes a deliberately vague public error message acceptable: the detail still exists, it just lives in logs. Two cautions. First, treat a caller-supplied id as untrusted input: validate its length and character set before logging it, or you enable log injection and unbounded cardinality. Second, keep it opaque and non-sensitive — never derive it from a user id or session token. In practice this is usually the trace id from W3C `traceparent`, so the same value joins logs, traces and support tickets.

code

http · 9 lines
http
HTTP/1.1 503 Service Unavailable
Content-Type: application/json
X-Request-Id: 4d5f1a0c9b2e47c8

{
  "code": "upstream_unavailable",
  "message": "The service is temporarily unavailable. Quote the request id when contacting support.",
  "requestId": "4d5f1a0c9b2e47c8"
}

go deeper

for a junior

Know that the id ties the client-visible error to the server logs and that the client quotes it to support.

for a middle

Describe the full lifecycle — assign at the edge, request-scoped logging context, echo in header and body, propagate downstream.

for a senior

Add the untrusted-input handling for inbound ids, metric cardinality risk, and verifying end-to-end that the id actually resolves in log search.

for a principal

Treat it as observability standardisation across services: one identifier reused from Trace Context, mandated by shared middleware, with a documented support workflow.

## The problem being solved A customer says "your API returned an error around 2pm". Without a shared key, finding that request means guessing from timestamps across several services. A **correlation id** (request id, trace id) is a single opaque value that identifies one logical request end-to-end so a support engineer can retrieve exactly what happened. ## Lifecycle **1. Assign at the edge.** The first component that sees the request — gateway, load balancer, or the service's own inbound filter — either accepts an incoming id or generates one. Generating it at the outermost hop means even requests rejected early are traceable. **2. Put it in request-scoped context.** Store it where the logging framework picks it up automatically (an MDC, an async-local, a context object). Manual passing to every log call is unreliable; automatic injection means *every* line for that request is tagged, including ones written by libraries. **3. Echo it to the client.** Two places, for two reasons: - A **response header** on *every* response, success or failure. Clients can log it for their own correlation, and it works even when the body is empty or the response comes from a proxy. - Inside the **error body**, because that is what a human copies out of a screenshot or a UI error panel. A UI should display it. **4. Propagate downstream.** Forward it on outbound HTTP calls and message publishes, so the whole call graph shares one key. Otherwise correlation stops at the first hop. ## Choosing the value Use a collision-resistant opaque identifier — a UUID, a ULID, or the 16-byte trace id from W3C Trace Context. Requirements: - **Opaque.** It must not encode a user id, tenant, email or token. It appears in client logs, screenshots and support tickets, all lower-trust than your database. - **Unique enough** that a search returns one request. - **Readable aloud/copyable.** Hex or base32 beats anything with punctuation a user might mangle. Because modern stacks already propagate `traceparent` for distributed tracing, the pragmatic choice is to surface the existing trace id rather than invent a parallel one — then a ticket id also opens the distributed trace. ## Accepting caller-supplied ids Accepting an inbound id is valuable — a client can correlate its own logs with yours — but the value is now untrusted input: - **Validate** length and charset (e.g. up to 64 hex/alphanumeric characters) and regenerate if it doesn't match. Otherwise you take newlines and control characters into your log stream, enabling log-injection and forged entries. - **Watch cardinality.** If this id becomes a metrics label or an index dimension, a client sending random values per request can blow up storage. Ids belong in logs and traces, not in metric labels. - **Never trust it for authorization or dedup.** It identifies a request for humans; it is not an idempotency mechanism. ## Operational payoff The payoff is only realised if the whole loop exists: id in the body, id in the logs, logs searchable by the id, and a documented support flow that asks for it. Teams frequently ship step one and skip the rest, producing an id that looks reassuring but resolves to nothing. Test it deliberately: trigger a failure, take the id from the response, and confirm one search retrieves the full server-side story.

  • Should the correlation id appear on successful responses too, or only on errors?
    Put it on every response via a header. Clients can then record it alongside their own request log, which makes it available later for issues that only surface after the fact — a slow call, a wrong-but-successful result. The body copy is specifically for errors, because that is the value a human copies out of a UI or screenshot.
  • How does the correlation id relate to a distributed trace id?
    They serve the same purpose at different granularities: a trace id spans the whole distributed call graph, while a span id identifies one hop. The clean approach is to surface the W3C Trace Context trace id as the correlation id, so one value joins logs, traces and support tickets instead of maintaining two parallel identifiers.

saying these in an interview costs you the question

  • Returning an id in the body that is never written to the logs, so it resolves to nothing
  • Deriving the id from a user id, email or session token instead of using an opaque random value
  • Logging a caller-supplied id without validating length or characters, enabling log injection
  • Only including the id on errors, so successful-but-wrong requests cannot be traced
  • Using the correlation id as an idempotency key or for authorization decisions

context