skip to content

Error Design

How a REST API communicates failure: machine-readable error bodies, the problem+json standard, validation-error shapes, and retry signalling. Interviewers probe it because error contracts are where API consistency and client resilience are actually won or lost.

part ofAPI stylesoverview, primer and where to startread it →
on this pageshow

questions

16

Why is it a problem for an HTTP API to return stack traces, SQL fragments, framework class names or internal database identifiers in an error response body, and what do you return instead?

level: juniorimportance: must knowfreq 62%

answer

  1. trace = free recon: version, packages, schema
  2. SQL text invites injection probing
  3. sequential IDs leak volume + enumeration
  4. generic message + correlation id
  5. global handler; test that no trace escapes

basics

~20 s

Those details leak your internals to attackers — stack, framework versions, table names, ID ranges — and are useless to callers. Return a stable code, a safe message and a request id; keep the trace server-side in logs, keyed by that id.

solid answer

~50 s

Two reasons. **Security**: a stack trace or SQL fragment reveals your framework and version, package layout, table and column names, file paths, and sometimes credentials in a connection string. That is free reconnaissance and turns a generic probe into a targeted attack. Internal identifiers are worse when they are sequential — they expose row counts and let an attacker enumerate resources. **Usefulness**: the caller cannot act on `NullPointerException at OrderService.java:214`. It tells them nothing about what to fix in their request. The fix is to split the audiences. The client gets a stable code, a safe human message, and a correlation id. The server logs the full exception, stack and context under that same id. Support then joins them. Enforce it with a global exception handler that maps unmapped exceptions to a generic `500` body, so no framework default trace page ever escapes — and check that the non-production behaviour isn't accidentally enabled in production.

code

http · 11 lines
http
HTTP/1.1 500 Internal Server Error
Content-Type: application/json

{"error":"org.postgresql.util.PSQLException: ERROR: column o.tenant_id does not exist\n\tat com.acme.billing.LedgerRepository.load(LedgerRepository.java:214)"}

--- instead ---

HTTP/1.1 500 Internal Server Error
Content-Type: application/json

{"code":"internal_error","message":"An unexpected error occurred.","requestId":"01J8YQ2R4K7ZC3M9"}

go deeper

for a junior

Say clearly that traces and SQL leak internals and help attackers, and that the client gets a generic message plus a request id instead.

for a middle

Add the concrete leak categories (framework version, schema, paths) and the global exception handler that guarantees the envelope.

for a senior

Cover how it leaks in practice — default error pages, ORM exception messages, gateway pages — and the regression test plus production config verification.

for a principal

Position it as an information-disclosure control in the threat model, define the org-wide rule for what may cross the boundary, and make it enforced by shared middleware rather than per-team discipline.

## What leaks, and why it matters An unhandled exception rendered into an HTTP response is one of the most common information-disclosure bugs in web APIs. What escapes is not just noise: - **Framework and version** (`Spring Boot 3.2.1`, `Express`, `Django`) — an attacker can look up known CVEs for that exact version. - **Package and file layout** (`com.acme.billing.internal.LedgerRepository`, `/srv/app/src/...`) — reveals architecture and, with source-path disclosure, sometimes lets an attacker guess unlisted endpoints. - **SQL text and schema** — a leaked query names tables and columns and shows whether the query is parameterised; it is the single most useful artifact for someone probing for SQL injection. - **Connection strings and hostnames** — internal DNS names, ports, occasionally credentials. - **Internal identifiers** — sequential primary keys disclose volume (order 10042 tells you roughly how many orders exist) and invite enumeration of other users' resources when authorization is weak. Even a *timing or wording* difference leaks: "user not found" versus "wrong password" turns a login endpoint into a user-enumeration oracle. The general rule is that an error body should reveal only what the caller is already entitled to know. ## Why it doesn't help the caller either A stack trace is a map of *your* code. The caller cannot change your code. What they need is: was this my fault or yours, what exactly in my request was wrong, is it worth retrying, and what do I quote to support. None of that is in a trace. ## The replacement design Split the audiences at the boundary: - **To the client:** a stable error code, a human-safe message, and a **correlation id** (request id / trace id). For a `500`, the message should be deliberately vague — "An unexpected error occurred" — because by definition you don't know what happened and can't safely characterise it. - **To your logs:** the full exception, stack, sanitized request context, user/tenant, and the *same* correlation id, so support can retrieve the detail in one query. That correlation id is what makes the vague public message acceptable: nothing is lost, it is just moved behind an authenticated boundary. ## How it actually escapes in practice Rarely by deliberate design — usually through defaults: - A framework's development error page left enabled in production (a common misconfiguration when a profile or env var is wrong). - A `catch` block that does `message = ex.toString()` or `ex.getMessage()` and puts it in the body. Driver and ORM exception messages routinely embed SQL and values. - Validation libraries echoing the rejected value back, which can reflect secrets the client sent. - Reverse proxies or gateways returning their own verbose upstream error pages. ## Enforcing it Use a single global exception handler that produces the standard error envelope for every unmapped exception, so no path renders a default trace. Have it map *known* exceptions to specific codes and everything else to a generic `500`. Then test it: an integration test that triggers an unmapped exception and asserts the body contains no stack marker (`at `, `Exception`, `select `) is cheap and catches regressions. Also verify the production configuration explicitly rather than assuming the default, and check gateway-level error pages, which sit outside your application code.

  • If the stack trace is hidden, how does a developer integrating with your API debug a 500?
    They quote the correlation id returned in the body and you look up the full trace in your logs. For self-service, expose the detail through an authenticated developer dashboard keyed by that id. The key point is that the detail still exists — it just lives behind an authentication boundary instead of in an anonymous HTTP response.
  • Is it acceptable to return full traces when a debug flag or non-production profile is enabled?
    Only if the flag cannot be turned on in production and is not client-controllable. A header or query parameter that switches on verbose errors is an attacker-controlled switch and should not exist. Environment-scoped configuration is safer, but it must be verified in production rather than assumed, since a misapplied profile is one of the most common ways traces leak.

saying these in an interview costs you the question

  • Assuming a stack trace is harmless because it contains no passwords
  • Putting ex.getMessage() straight into the response body, which often embeds SQL and parameter values
  • Relying on a debug flag that a client can flip via a header or query parameter
  • Believing internal numeric IDs are safe to expose because they are meaningless — sequential IDs leak volume and enable enumeration
  • Returning different messages for unknown user versus wrong password, creating an enumeration oracle

context

open as a page

What is the application/problem+json media type defined by RFC 9457, and which members does a problem document define?

level: juniorimportance: must knowfreq 47%

basics

~20 s

It is a standard JSON format for HTTP error bodies, media type application/problem+json. Its members are type (a URI identifying the problem kind), title, status, detail and instance — all optional, plus your own extension members.

open as a page

What does the HTTP Retry-After response header mean, what value formats does it accept, and on which responses should an API send it?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Retry-After tells the client how long to wait before trying again. It takes either delay seconds (Retry-After: 120) or an HTTP-date. Send it on 429 and 503, and on 3xx redirects that ask the client to wait.

open as a page

A client posts a JSON body to your REST API and several fields fail validation. How would you shape the response body so the caller can highlight the exact inputs that failed, rather than returning one human-readable sentence?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Return an array of violations, one per failed field. Each entry carries a machine-readable code, a target naming the field, and a human message. Clients switch on the code and attach the message to the right input; a prose string cannot be parsed.

open as a page

What should the body of an HTTP API error response contain, and why do teams separate a stable machine-readable error code from the human-readable message?

level: middleimportance: must knowfreq 68%

basics

~20 s

An error body should carry a stable machine-readable code, a human message, and a request id for support. Clients branch on the code; the message is prose that can be reworded or localized without breaking any caller.

open as a page

A request body is well-formed JSON and matches the declared content type, but breaks a schema or business rule. Would you answer with HTTP status 400 or HTTP status 422, and what distinction do those two codes actually carry?

level: middleimportance: must knowfreq 68%

basics

~20 s

400 means the request itself is malformed - bad syntax, unparseable JSON, missing required parameter. 422 means the syntax parsed fine but the content was semantically unprocessable. Both are non-retryable client errors; pick one convention and apply it API-wide.

open as a page

How do you decide which errors in an HTTP API contract are worth retrying and which are permanent, and how do you communicate that distinction to client developers?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Retryable means the same request could succeed later without changing: 429, 503, 504, connection failures. Permanent means the request itself is wrong: 400, 401, 403, 404, 422. Mark retryability explicitly per error code in the contract, and require idempotency before retrying writes.

open as a page

How do you use a correlation or request id in HTTP API error responses so that support can trace a reported failure, and where should that id come from?

level: middleimportance: should knowfreq 48%

basics

~20 s

Generate or accept one id per inbound request, put it in every log line and in every error body (and ideally a response header), and propagate it to downstream calls. Support then searches logs by the id the caller quotes.

open as a page

In an RFC 9457 problem details response, what is the type member for, and how should you choose its URI — does it have to be a URL that resolves?

level: middleimportance: should knowfreq 36%

basics

~20 s

The type member is a URI identifying the kind of problem — the stable key clients branch on. It need not be dereferenceable, but a resolvable HTTPS documentation URL is recommended. Clients must compare it as an opaque string, never fetch it at runtime.

open as a page

What do the RateLimit response headers (and the older X-RateLimit-* convention) tell a client, and how do they relate to HTTP 429 and Retry-After?

level: middleimportance: should knowfreq 40%

basics

~20 s

They advertise the quota state before you hit the wall: the limit, how much remains, and when the window resets. Retry-After tells you what to do after a 429; rate-limit headers let a client pace itself and never get one.

open as a page

When a submitted request breaks a dozen rules at once, do you return the first failure or all of them in one response? What are the consequences of each choice for the caller and for the server?

level: middleimportance: should knowfreq 52%

basics

~20 s

Aggregate: run all cheap validators and return every violation in one array, so a form can be fixed in one pass instead of one round trip per mistake. Fail fast only across layers - parse errors stop everything, and expensive or stateful checks run only after cheap ones pass.

open as a page

How do you add API-specific data to an RFC 9457 problem details response, and what rules govern those extension members so the contract can evolve safely?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Add them as extra top-level members of the problem JSON object — no nested envelope needed. Consumers must ignore unrecognised members, so adding one is non-breaking; renaming, removing or changing a member's type is breaking.

open as a page

In a validation error response, how do you identify the offending value when it sits deep inside the payload - say the quantity of the fourth element of an items array? What does an RFC 6901 JSON Pointer give you, and where does it fall short?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Use a JSON Pointer into the request document: /items/3/quantity. It is a slash-separated path of member names and zero-based array indexes, empty string for the whole document, with ~1 escaping a literal slash and ~0 a literal tilde. It cannot address query parameters or headers.

open as a page

How do you keep error response bodies consistent across every endpoint of an HTTP API — and across many services in one organisation — and what concretely breaks when they are inconsistent?

level: principalimportance: should knowfreq 38%

basics

~20 s

Enforce one envelope in shared middleware, not by convention: a single global exception handler plus a shared library or gateway normalisation, with contract tests. Otherwise every client writes N parsers and error handling silently rots at the edges.

open as a page

Your service returns HTTP 503 with a correct Retry-After header during a partial outage, clients obey it, and the service still collapses when it comes back. What retry behaviour do you require from clients and what do you enforce server-side?

level: principalimportance: should knowfreq 38%

basics

~20 s

Obeying a single Retry-After synchronises every client into one burst at recovery. Require exponential backoff with jitter, capped attempts, an overall deadline and circuit breakers; server-side, randomise the advertised delay, shed load, and ramp capacity back gradually.

open as a page

How do server frameworks such as Spring Boot and ASP.NET Core produce application/problem+json responses out of the box, and what do you typically have to customise?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

Both ship first-class support: Spring Boot has a ProblemDetail type and an opt-in switch to return problem+json for framework errors; ASP.NET Core returns problem documents by default for error statuses and validation failures. You customise the type URIs, extensions and detail text.

open as a page