skip to content

How would you make failures that happen after a response is committed visible in production, given they never reach the error mapper?

level: seniorimportance: should knowfreq 44%

answer

  1. the recorded status is the frozen one
  2. success dashboards cannot see it
  3. needs its own counter
  4. mark the record as post-commit
  5. correlation identifier in a header before flush

basics

~20 s

Stop relying on response status: the committed status is what gets recorded, so a half-sent success logs as a success. Emit a dedicated counter, mark the record post-commit, and set a correlation header before the first flush.

solid answer

~50 s

These failures hide because every default observability signal is derived from the **status that was committed**. The access record and every alert built on server-error ratio read a `200` frozen before the failure existed, so the incident is real for users and invisible to the on-call. Three things fix it. First, the write path must **record the failure explicitly** and mark it as post-commit, with the request's correlation identifier and how much of the body had been emitted. Second, the class needs **its own counter and alert**, independent of mapped error responses and separate from streams the peer stopped reading. Third, the correlation identifier has to be written into a **response header before the first flush**, because after commit you can no longer give the user anything to quote back to support. Without that, a truncated download and a server-side record cannot be joined.

go deeper

for a junior

Understand that the status written into the access record is the one the server committed, not a verdict on whether the whole body arrived. A request logged as successful can still have delivered half an answer.

for a middle

Explain why the standard error-rate signals cannot see this: they are derived from sent statuses. Describe the explicit record the write path has to emit instead, and what it should contain.

for a senior

Show the operational wiring: a dedicated counter labelled by endpoint, an alert on a normally-zero rate, causes kept separate, and a correlation identifier emitted as a header before the first flush so support can join a user report.

for a principal

Treat it as a platform gap rather than one service's bug. The write path that streams is shared, so the post-commit signal should be emitted there once and required of any endpoint permitted to stream.

## Why this class of failure hides Every standard view of a request is keyed on its status code, and after commit that code is frozen at whatever optimistic value was chosen before the failure existed. The consequences compound: - The **access record** shows a success, so success-rate dashboards are unmoved. - The **server-error ratio alert** never fires, because no server error was ever returned. - The **latency view** may even look better, since an aborted response finishes sooner than a complete one. - **Client-side** telemetry often agrees, because a cleanly terminated truncation looks like a normal response to the client too. The result is the worst incident shape there is: users see wrong or missing data, and every internal signal says the service is healthy. Somebody has to notice by hand. ## What to record at the moment of failure The write path is the only place that knows what happened, so it must say so explicitly rather than letting the exception die where it was caught: 1. **The failure itself**, with the exception and the operation that was in flight, captured at the point where the write loop gave up. 2. **An explicit post-commit marker** — a structured field saying the response had already been committed. Without it, a reader sees an exception on a request logged as `200` and assumes it was handled. 3. **The request's correlation identifier**, so a server-side record can be joined to whatever the user reports. 4. **How much had been emitted** — bytes, records, or pages. This is what tells you whether the consumer got nothing or almost everything, which is the difference between an obvious outage and silent data drift. 5. **What the stream was producing** — the query, the export, the resource identifier — because the reproduction almost always depends on the particular data that failed. ## Wiring it to something that pages - Give post-commit failures **their own counter**, not a bucket inside general handler errors, and alert on it directly. A rate that is normally near zero makes a good alert; a ratio against total requests does not, because the denominator hides it. - **Separate the causes.** A failure your code raised and a response the peer stopped reading are different events with different owners and different urgency; counting them together produces an alert nobody trusts. - **Label the counter by endpoint**, since streaming endpoints are few and knowing which one broke is usually most of the diagnosis. - **Do not** build the alert on connection resets seen by the network layer alone. Those are noisy for reasons that have nothing to do with your handler, and they miss the case where the framework finished the message cleanly. ## Which signal sees what | Signal | Sees a mapped error response | Sees a post-commit failure | |---|---|---| | Server-error ratio from sent statuses | Yes | No — the sent status was a success | | Access record for the request | Yes, with the error status | Only as a successful request | | Client-side success telemetry | Yes | Not when the message ended cleanly | | Dedicated post-commit counter | Not its job | Yes — this is the only direct signal | | Exception record from the write path | Yes | Yes, if the write path emits one | The table is the whole argument for a dedicated signal: for this class, four of the five rows a service normally relies on report nothing, and the two that work are ones somebody has to add on purpose. ## Giving support something to join on After commit you cannot put anything into headers, so anything the user might quote back has to be emitted **before the first flush**. In practice that means the correlation identifier is a response header set by the layer that assigns it, not something written into a failure body — because for this class there is no failure body. The same identifier belongs in whatever the streamed payload's contract offers up front, if it has one, so a saved file still carries the thread back to the server record. ## Separating the two causes that look alike A stream can end early for two very different reasons: the server failed while producing it, or the peer stopped reading. They arrive at the same place in the code and are easy to record as one event, which is a mistake — one is a defect owned by the service, the other is routine on endpoints that people cancel. Counting them together produces a number that is always non-zero and therefore never alerted on, which is how a real signal gets buried inside a noisy one. Keep them as separate counters even when the handling code path is shared. ## The cheap structural test A useful exercise on any service that streams: take one endpoint, force a failure partway through the body, and then ask what a stranger could learn from the outside. If the answer is *the request looks successful and no signal changed anywhere*, the observability for this class does not exist yet, whatever the general error dashboards say. Fixing it is a small amount of code in one place — the write path — and it is the difference between hearing about half-sent responses from your own alerting and hearing about them from a customer weeks later.

  • Why is a server-error-ratio alert useless for this class of failure?
    Because the ratio is computed from the statuses that were sent, and the status here is a success that was committed before anything failed. The failing request lands in the numerator of the healthy side, so the ratio can stay flat through a complete outage of a streaming endpoint.
  • Why must the correlation identifier travel in a response header rather than only in a failure body?
    Because for a post-commit failure there is no failure body to put it in — the body is whatever partial content already went out. The header must be written before the first flush; once the response is committed, no header can be added.
  • What should be recorded besides the exception itself?
    An explicit post-commit marker, the correlation identifier, how many bytes or records were emitted before the failure, and which resource was being produced. The emitted count is what separates a consumer that got nothing from one that silently stored an almost-complete result.

saying these in an interview costs you the question

  • Assumes the existing error-rate dashboard already covers these failures
  • Reads the recorded status as proof the client got a complete response
  • Counts post-commit failures in the same bucket as mapped error responses
  • Puts the correlation identifier only in the failure body, which never exists here
  • Relies on client complaints as the detection mechanism
  • Alerts on raw connection resets and calls the class covered