skip to content

How would you decide which routes get a replayable request body across a platform, and what does that cost?

level: principalimportance: nice to knowfreq 35%

answer

  1. streaming stays the default
  2. capture is opt-in per route
  3. concurrency times cap equals memory
  4. one pass beats a copy
  5. loud rejection, never silent truncation

basics

~20 s

Keep streaming as the default and make capture an opt-in per route, with a measured cap and a loud failure past it. Worst-case memory is peak concurrency times that cap, so upload and relay routes stay streaming permanently.

solid answer

~50 s

Treat replayability as a capacity decision, not a convenience. The default stays **stream once**, because a captured payload costs memory equal to peak in-flight requests multiplied by the cap, and the caller picks the size. Then ask who genuinely needs a second read: verifying a signature over the raw bytes, auditing a small number of high-value operations, deduplicating by payload digest, or sampling for a live investigation. Most other cases are two components that each wanted the *parsed* value, which is solved by sharing the result rather than copying the bytes. Where capture is justified, give it an explicit cap derived from measured body sizes, reject with `413` past it rather than silently truncating, exempt upload and relay routes by name, and ship one shared implementation so nobody writes an unbounded version of their own.

go deeper

for a junior

Take away the default: payloads stream and are read once, and anything that wants a second look has to arrange it deliberately rather than assuming it is free.

for a middle

Be able to state the cost in one line — memory equals concurrent requests times the cap — and name the alternative of doing the extra work during the single pass.

for a senior

Argue route by route: which endpoints genuinely need a second read, what the cap is, what happens when it is hit, and which routes are exempt because their payloads are large.

for a principal

Make it a platform position with one shared implementation, measured caps, a loud failure mode, a named exemption list, and a review trigger when traffic shape or a hop changes.

## Start from the default: stream Payloads are read once because keeping them costs memory that scales with concurrency times body size, and the caller chooses the size. So the platform default is **streaming, single reader**, and replayability is a feature a route asks for and pays for — never a global convenience switched on because one team hit an empty-body bug. Framing it that way turns a debugging annoyance into a capacity decision, which is what makes this a platform-level question rather than a per-service one. ## Who genuinely needs a second read Ask what the second reader is for. In practice the honest cases are few: - **Verifying a signature over the raw bytes** before anything parses them, then parsing them again. - **Auditing the exact payload** for a small number of regulated or high-value operations. - **Deduplicating by payload digest** for idempotent submissions. - **Capturing a sample** for diagnostics on a route that is actively being investigated. Everything else usually turns out to be two components that each wanted the parsed value, which is a structural problem, not a buffering one. ## Cheaper answers that avoid a copy - **Do the work in the single pass.** A digest can be computed incrementally as the bytes stream through, so verification needs no copy at all. Re-serialising a parsed object to recover the original bytes does not work — formatting, ordering and whitespace differ, and the digest will not match. - **Share the parsed result, not the stream.** One component reads and parses; later stages consume that value. - **Capture a bounded prefix.** For diagnostics, the first few kilobytes usually answer the question at a fixed, tiny cost. - **Sample.** Capture on a small fraction of requests, or only while an investigation is open. ## The cost model to put on the table The number that decides the policy is simple: > worst-case additional memory = peak in-flight requests on a route x the cap Set the cap from measurement, not from taste: look at the body-size distribution per route and pick a limit above the real traffic and far below what the process can survive at peak concurrency. Then state what happens at the limit: | Behaviour at the cap | Good for | Bad for | |---|---|---| | Reject with `413` | anything that must see the whole payload | routes with legitimately large bodies | | Capture a prefix, mark truncated | diagnostics and sampling | signature checks and audit | | Silently stop capturing | nothing | everything — it reintroduces the empty-body bug | The last row is the one to rule out explicitly. A policy whose failure mode is silence recreates exactly the class of bug that motivated the feature. ## Who decides, and on what evidence A route-by-route policy only works if someone can answer "why is this route capturing?" a year later. Three things make that possible: - the **reason** recorded next to the route, in the same place as the cap, so a reviewer sees both; - a **named owner** for the exemption list, because the expensive mistakes are additions nobody challenged; - a **default answer of no**, so the burden of proof sits with the request rather than with the platform team refusing it. ## Exemptions that must be permanent Upload routes, routes that relay a payload onward, and anything with a realistic body in the megabytes stay streaming, permanently and by name. If they are exempted only by falling under a cap, someone will raise the cap. Spill-to-disk is the middle option: it removes the memory ceiling and adds latency, disk pressure, and a cleanup obligation on every exit path including errors and timeouts. It buys real capacity, and it is worth it only where the audit or verification requirement is genuine. ## Making it a decision rather than a habit 1. Default streaming; capture is opt-in per route, with the cap and the reason recorded next to the route. 2. Provide one shared implementation of the capture wrapper so teams are not each writing a different unbounded version. 3. Measure it: body-size percentiles per route, count of cap rejections, memory attributable to capture. A cap with no alarm is a cap nobody will notice hitting. 4. Review the exemption list when traffic shape changes, not when an incident forces it. ## What a strong answer sounds like It refuses the global switch, names the small set of legitimate second readers, proposes the single-pass alternative first, puts an explicit number on the memory cost, chooses a loud failure at the cap, and keeps large-payload routes permanently out of scope.

  • A team wants to verify a signature over the raw payload. How do you satisfy them without a copy?
    Compute the digest incrementally while the single reader streams the bytes through, then let binding parse the same pass. What does not work is re-serialising the parsed object to recreate the bytes: formatting, ordering and whitespace differ from what arrived, so the digest will not match and the failure looks intermittent.
  • How do you choose the cap, and what tells you it is wrong?
    From measurement: per-route body-size percentiles, a limit comfortably above real traffic and far below what the process survives at peak concurrency. Alarm on rejections at the cap and on memory attributable to capture. Frequent rejections mean the route was mis-classified; memory tracking traffic means the cap or the exemption list needs revisiting.

saying these in an interview costs you the question

  • Turns capture on globally because one team hit an empty-body bug
  • Sets no cap, or raises it whenever a large request is rejected
  • Silently stops capturing at the limit, recreating the empty-body bug
  • Re-serialises a parsed object to recover the original bytes for a digest
  • Captures upload and relay routes alongside ordinary small-payload routes
  • Spills to disk without cleaning up on error and timeout paths