A gRPC server-side interceptor catches every handler exception so its audit write always happens; why do callers now see OK (0)?
answer
- the wrapper is in the path, not beside it
- catching is fine, not re-raising is not
- completing normally writes OK (0)
- client and server dashboards agree, both wrong
- record the outcome, propagate it unchanged
basics
~20 sBecause the wrapper sits between the handler and the machinery that writes the call's status. Catching the failure and then completing the call normally means OK (0) goes on the wire, and the failure exists only in the server's own log.
solid answer
~40 sA server-side interceptor is the last thing to touch a call before its trailing status is written. Catching the handler's error and then completing the call normally does not merely *observe* the failure — it *replaces* it. The caller receives `OK (0)`, indistinguishable from a real success, so its error counters stay at zero, its retry logic never fires, and any wrapper nested outside this one counts a success too. The failure now exists in exactly one place: this server's log. The fix is to record the outcome and then let the original failure propagate unchanged, so the audit write and the caller's status both happen. The mirror of this bug lives on the client side, where a wrapper that returns a default value on failure hides the same thing from the calling code.
code
pseudocode · 13 lines# the defect
function auditInterceptor(call, request, nextHandler):
outcome = attempt nextHandler(call, request)
if outcome.failed:
writeAuditRecord(call.method, outcome.error)
return completeNormally(emptyResponse) # caller now sees OK (0)
return outcome
# the fix
function auditInterceptor(call, request, nextHandler):
outcome = attempt nextHandler(call, request)
writeAuditRecord(call.method, outcome.status) # success or failure, recorded
return outcome # failure forwarded unchangedgo deeper
Remember that a server-side wrapper sits in the path of the call, so returning normally after catching an error means the caller is told the call succeeded.
Explain the mechanism: the outcome the wrapper returns becomes the call's trailing status, so completing normally writes OK (0) and the client's success branch is taken.
Show the blast radius — client counters, retry policy and the server's own outer metrics all report success — and give the fix: record on every path, forward the original outcome unchanged.
Treat it as a contract question. If a shared wrapper is mandatory across the estate, its ability to rewrite an outcome is estate-wide risk, and the review bar for that code is not the same as for a service handler.
## Where the wrapper sits, and why that is the danger On the server side an interceptor is not an observer bolted to the side of the call. It is in the path: the outcome the wrapper returns *is* the outcome the machinery turns into the call's trailing `grpc-status`. Anything the wrapper does to that outcome is what the caller gets. So there are two very different-looking pieces of code with identical intent and opposite effects: - **Record and forward** — write the audit entry, return the outcome unchanged. Caller sees the real status. - **Record and complete** — write the audit entry, then finish the call normally. Caller sees `OK (0)`. The second is usually written for a good reason: somebody wanted the audit write to happen *even when the handler blows up*, and reached for a construct that catches everything. The catch was the right instinct; not re-raising was the defect. ## What the caller actually experiences Nothing. That is the point. A call that ends `OK (0)` is a successful call by every definition the client has: - the client library returns normally, so the calling code takes its success branch; - client-side outcome counters increment the success bucket; - any client-side retry policy keyed on a failure status has nothing to act on; - a client-side wrapper measuring outcomes records a success as well. And because the status travels in **trailing** metadata, there is nothing earlier in the exchange that contradicts it — there is no separate signal a careful client could have consulted instead. ## The blast radius inside the server The server is not spared either. Post-work unwinds inside out, so every wrapper nested **outside** the swallowing one observes the rewritten outcome. A metrics wrapper placed outermost — the usual advice — now reports a healthy service. The service's own dashboards agree with the caller's dashboards, and both are wrong in the same direction. In a court records e-filing service that combination is the worst case available: the clerk's client reports every filing accepted, the server's success rate reads 100%, and the filings did not persist. The only surviving evidence is a log line nobody is alerted on. ## Doing it correctly 1. **Record the outcome, then propagate it.** Whatever construct guarantees the audit write, it must end by re-raising or returning the original failure rather than a fresh success. 2. **Record failures as failures.** The audit entry carries the terminal status, not just the fact that a call happened — an audit trail whose entries all read "completed" is the same defect one layer down. 3. **Do not invent a status.** A wrapper's job is to report what happened; deciding that a missing case file is `NOT_FOUND (5)` rather than `INTERNAL (13)` is the handler's call, made with knowledge the wrapper does not have. 4. **Let the unmapped case stay loud.** An exception the server does not map surfaces as `UNKNOWN (2)`. That is noisy and that is correct — it means somebody must decide what the right status is, and a swallowing wrapper silences exactly that signal. ## Distinguishing the two failure modes in an interview They are often conflated, and naming both is what separates a good answer: | | wrapper swallows the failure | handler raises an unmapped error | |---|---|---| | status the caller sees | `OK (0)` | `UNKNOWN (2)` | | client's success branch | taken | not taken | | retry policy | never triggers | may trigger | | how it is found | someone notices missing data | it shows up in error metrics immediately | | severity | silent corruption of the contract | noisy, but honest | The second is a bug. The first is a bug that also disables the mechanisms that would have found it. ## The client-side mirror The same shape appears on the calling side. A client-side wrapper that catches a failed call and returns an empty or default response spares the calling code an error path and removes the caller's own ability to react. It is rarer, because a wrapper on that side usually has no plausible default to return — but where one exists (an empty list, a cached value), it is the same defect with the same silence. ## The rule worth stating plainly A wrapper may **observe** an outcome, **annotate** it with metadata, or **map** it deliberately as an explicit policy. What it must not do is **convert a failure into a success as a side effect of trying to log it** — and that is the single most common way a well-intentioned audit wrapper breaks a service.
- Why do the server's own metrics agree with the caller that nothing failed?Post-work unwinds inside out, so a metrics wrapper nested outside the swallowing one observes the rewritten outcome, not the handler's. Both sides therefore count a success. Only a wrapper nested *inside* the swallowing one would still see the truth, which is not where metrics usually sit.
- Should the interceptor pick a better status than UNKNOWN (2) when the handler raises something unmapped?No. The wrapper knows the method name, not why the call failed, so any status it picks is a guess dressed as a fact. `UNKNOWN (2)` is noisy on purpose: it marks a case somebody needs to map deliberately in the handler, where the knowledge lives.
- How do you guarantee the audit write without touching the caller's status?Record the outcome on every path — success and failure — and return the original outcome untouched. The write and the propagation are independent; the bug comes from using the construct that guarantees the write to also produce the return value.
- Is there a legitimate reason for a wrapper to change a status?Yes, as an explicit policy rather than a side effect: mapping an internal failure to a coarser status at a trust boundary so the detail does not leak outward. The difference is that it is deliberate, documented, and never turns a failure into a success.
saying these in an interview costs you the question
- Thinks logging an error is the same as propagating it
- Says the client can still detect the failure some other way
- Believes the server's metrics would catch it
- Has the wrapper choose a status the handler did not produce
- Treats UNKNOWN (2) as something a wrapper should quietly hide
- Assumes a client-side wrapper cannot make the same mistake