skip to content

You inherit an HTTP API where almost every endpoint is a POST to a verb-shaped path returning 200 with a status field in the body. How do you judge whether that is a real problem, and what would you do about it?

level: principalimportance: should knowfreq 34%

answer

  1. 200-with-error-body blinds gateways and SLOs
  2. POST reads = no CDN, no ETag
  3. no URL = no linking from tickets/logs
  4. timeouts → duplicates without idempotency keys
  5. fix status codes first, or go honest RPC

basics

~20 s

Judge by cost, not purity: are resources unaddressable, are errors invisible to monitoring and gateways, are reads uncacheable, do retries misbehave? If yes, fix incrementally — model the core entities as addressable resources, use real status codes — or adopt an explicit RPC framework rather than a half-hearted hybrid.

solid answer

~60 s

First I separate style from damage. An API being POST-heavy is not automatically wrong — operation-centric domains exist. The questions I ask are operational: - **Addressability**: can a support engineer fetch an entity by URL? If everything is `POST /getThing`, no. - **Observability**: do failures show as 4xx/5xx? `200` with `{"status":"error"}` makes gateways, dashboards, and SLOs blind to real failures — usually the most expensive symptom. - **Caching**: are hot reads POSTs? Then no cache, CDN, or conditional request helps. - **Retries**: with no idempotency story, timeouts turn into duplicates. - **Cost of change**: how many clients, and can they be redeployed? Then I pick a direction. If the domain really is procedural, the honest move is an explicit RPC surface (gRPC, or documented JSON-RPC) rather than pretending. If it is entity-centric, I migrate incrementally: expose the core nouns as `GET`-able resources, return real status codes, add an idempotency-key contract on writes, and route new work to the new shape. What I avoid is a purity rewrite with no measurable benefit.

go deeper

for a junior

Recognise the pattern and name one concrete downside, such as errors hidden inside a 200 response.

for a middle

List the operational costs — caching, status codes, addressability — and suggest incremental fixes rather than a rewrite.

for a senior

Prioritise fixes by value and risk, propose an idempotency contract for writes, and describe a coexistence plan for existing clients.

for a principal

Make the explicit build/keep/replace call with a cost argument, choose a coherent target (resource-oriented or honest RPC), sequence the migration by operational pain, and put linting and guidelines in place so the debt stops growing.

## Distinguish the smell from the harm "RPC over POST" describes an HTTP API whose paths are procedure names — `/getCustomer`, `/updateCustomer`, `/sendInvoiceEmail` — all invoked with POST, always answering `200 OK`, with success or failure indicated by a field in the JSON body. It is worth naming, but as a principal engineer your job is not to grade it against a style guide; it is to decide whether it costs enough to change and, if so, in which direction. ## The concrete costs **Errors are invisible to infrastructure.** This is usually the biggest one. Load balancers, API gateways, CDNs, tracing systems, and alerting all key on status codes. An API that returns `200` for a failed payment means your error-rate dashboards read zero during an outage, retries never trigger, and circuit breakers never open. Everything downstream has to parse bodies to learn the truth — and most of it will not. **Nothing is cacheable.** POST responses are effectively uncacheable in practice, so hot reads cannot use a CDN, a shared cache, `ETag`/`If-None-Match` revalidation, or browser caching. Teams then rebuild caching inside the application, badly. **No addressability.** With no stable URL per entity, you cannot link to a record from a ticket, a log line, or another system's payload. Debugging requires the right tool and the right body; `curl` is not enough. **Retry semantics are undefined.** POST is neither safe nor idempotent by default. Without an idempotency-key contract, every network timeout is a coin flip between a lost request and a duplicated one — a real financial problem in payments and ordering. **Authorization and rate limiting get harder.** Policies that would naturally be expressed by method and path prefix now require body inspection, so a gateway cannot enforce "reads are open, writes require elevated scope". ## What is *not* a real cost Not conforming to a style. If the API is internal, has three clients you deploy yourself, works, and none of the above symptoms bite, then "it is not RESTful" is not a business case. Rewriting a working surface for taxonomy is how platform teams lose credibility. Some domains — workflow engines, batch processors, ML inference — are genuinely operation-centric, and forcing resource nouns onto them produces contortions such as `POST /translations` for what everyone calls "translate". ## Choosing a direction **Direction A — go properly resource-oriented.** Appropriate when the domain really is entity-centric (customers, orders, invoices) and consumers are many or external. Incremental path: (1) expose the core entities as `GET /customers/{id}` alongside the existing endpoints, so reads become cacheable and linkable immediately; (2) start returning real status codes — this is often the highest value-per-effort change and can be done endpoint by endpoint; (3) introduce an idempotency-key header contract for writes; (4) reshape genuine actions as state-transition sub-resources or explicitly-marked custom methods; (5) freeze the old surface for new work and let it decay behind a documented sunset. **Direction B — commit to RPC honestly.** If the domain is procedural and the clients are internal, adopting gRPC or a documented JSON-RPC surface gives you generated clients, schema evolution, streaming, and clear semantics — everything the hand-rolled version approximates. Half-measures are the worst outcome: a surface that is neither browsable nor schema-driven. ## Sequencing and governance Whatever the direction, sequence by pain: fix status codes first (cheap, immediately improves observability), then addressability for the entities support staff ask about most, then idempotency on the money paths. Publish the target shape as a written guideline with linting so new endpoints stop adding to the debt, and accept a long coexistence — a hybrid where the old surface is stable and shrinking is a healthier state than an ambitious rewrite that stalls at 40%. The judgement to articulate in an interview is exactly this: name the concrete operational costs, weigh them against migration cost and client count, choose a coherent target rather than a purer one, and migrate along the axis that pays first.

  • Which single change would you make first in such an API, and why?
    Returning truthful HTTP status codes. It is usually a small, endpoint-local change, and it immediately restores error-rate dashboards, alerting, retry logic, and circuit breakers across every layer of infrastructure — none of which can read a JSON status field. Addressability and idempotency matter, but they cost more per endpoint.
  • How do you decide between migrating toward resources and adopting gRPC instead?
    By audience and domain shape. External or numerous consumers who want browsable, cacheable, curl-able endpoints argue for resources; internal service-to-service traffic in an operation-centric domain argues for gRPC, where you gain schemas, generated clients, and streaming. The failure mode is choosing neither and keeping an ad-hoc surface that has the benefits of both only in slide form.

saying these in an interview costs you the question

  • Calling for a full rewrite purely because the API is not RESTful, with no cost analysis
  • Missing that 200-with-error-body breaks gateways, alerting, and retries
  • Assuming POST responses can be CDN-cached with the right headers
  • Believing every domain can be modelled cleanly as CRUD resources
  • Proposing a big-bang cutover for an API with unknown external consumers

context