A webhook receiver caps bodies at 1 MB, yet its heap profile spikes on deeply nested JSON. Why?
answer
- bytes in, objects out
- two bytes of brackets buy a whole slice
- an any is where structure becomes allocation
- multiply by the requests in flight
- profile by allocation volume, not just live bytes
basics
~20 sA byte cap bounds input, not the decoded graph. Nested arrays decoded into an any allocate a fresh slice and interface value per level, so a megabyte of brackets becomes many megabytes of small heap objects — multiplied by every request in flight.
solid answer
~50 sThe cap does its job: no request reads more than 1 MB. But a megabyte of `[[[[...]]]]` decoded into an `any` produces one `[]any` plus its backing array and an interface value for every nesting level, so two bytes of input can buy tens of bytes of live heap and one small object each — hundreds of thousands of objects from one body, plus deep recursion in the decoder. A heap profile read by allocation volume points straight at the decode path, and the goroutine count tells you the multiplier: worst-case memory is roughly the byte cap's expansion factor times the requests in flight, not the byte cap. The fixes are structural: decode into a concrete struct so values you do not model are skipped rather than materialised, bound the depth yourself with a `json.Decoder` token loop, lower the byte cap for that route, and bound concurrency so the worst case is a number you chose rather than one the poster chooses.
code
go · 21 linesdec := json.NewDecoder(r.Body)
depth := 0
for {
t, err := dec.Token()
if err == io.EOF {
break
}
if err != nil {
return err
}
if d, ok := t.(json.Delim); ok {
if d == '[' || d == '{' {
depth++
if depth > 32 {
return errors.New("input nested too deeply")
}
} else {
depth--
}
}
}go deeper
Take away the core fact: the size of the JSON text and the size of the Go values it becomes are different numbers, and nested structure is where they diverge most.
Explain what a decoder allocates per nesting level when the target is an any, and why many small objects cost more than their bytes suggest.
Confirm it from a heap profile rather than asserting it, then give the peak as cap times expansion times concurrency and propose which term to bound first on this endpoint.
Own the posture: which endpoints may accept open-ended structure at all, what depth and concurrency bounds ship by default, and what evidence justifies relaxing them for one team.
## The shape of the surprise A public webhook receiver caps each request body with `http.MaxBytesReader` at 1 MB, and on-call still sees heap climb sharply under a burst from one unauthenticated poster. Nothing is leaking — the memory is released after the requests finish — but the peak is far above what "1 MB per request" suggests, and the capacity planner wants a number for what one request may cost. The reason is that a byte cap bounds the *encoded input*, and the memory is spent on the *decoded representation*. Those two are related by an expansion factor that the sender controls. ## Where the expansion comes from Decode a body of nested arrays into an `any` (or a `map[string]any`) and every level of nesting becomes real Go data: - a `[]any` slice header, - a backing array for that slice, with capacity that grows in steps as the level fills, - an interface value describing each element, - and, for objects, a `map[string]any` plus one string key per field. The input `[[[[[[…]]]]]]` spends two bytes per level. The decoded value spends a slice header, a backing array and an interface value per level. A body of `[{},{},{},…]` spends about four bytes per element and buys a map each. Either way, a megabyte of input becomes a very large number of *small* objects — and small objects are the expensive kind, because each carries allocator and garbage-collector bookkeeping and none of them are contiguous. There is a second cost on the same input: decoding a nested structure recurses, so nesting depth translates into goroutine stack growth for the goroutine serving that request. ## The multiplier nobody writes down The per-request number is only half of it. A server handles many requests at once, so the worst case is: peak ≈ (bytes per request) × (expansion factor) × (requests in flight) The first factor is the one you set. The second is chosen by whoever posts to you. The third is unbounded by default: `net/http` will start a goroutine per connection and your handler will happily decode in all of them. That product, not the 1 MB, is the number the capacity planner is asking for. ## Confirming it rather than guessing Take a heap profile and read it by allocation volume rather than by live bytes: allocation volume attributes the churn to the call sites that produced it, and the decode path will dominate. Live-bytes samples taken during a burst show the same peak from the other side. Cross-check with the number of goroutines serving requests at the peak to recover the multiplier. If the profile blames the decode path and the input is small, expansion is your answer; if it blames something after the decode, the payload is not the problem and you are chasing the wrong cap. ## What actually reduces it **Decode into a concrete struct.** This is the largest single win and it is often free. When the decoder is filling a typed struct, input that does not correspond to a modelled field is scanned past rather than materialised — the nested arrays under a field you never declared cost parsing time but not a graph of `[]any`. Decoding into `any` is what turns arbitrary structure into arbitrary allocation. **Bound depth explicitly** when you genuinely must accept open-ended shapes. A `json.Decoder` token loop lets you count opening and closing delimiters and reject beyond a depth you choose, which is a limit on structure rather than on bytes. **Lower the byte cap on that route.** A webhook payload with a known schema rarely needs a megabyte; a cap sized from the observed distribution with headroom shrinks every term in the product. **Bound concurrency.** A semaphore around the decode, a cap on in-flight requests, or a limit on accepted connections turns the third factor into a number you chose. Without it, no per-request cap gives you a peak. **Reject early where you can.** If the declared length is already over the cap, reject before reading; it costs nothing and saves the whole pipeline. ## The sentence to leave the reviewer with A size cap is necessary and not sufficient. It bounds what you read; it does not bound what that input expands into once decoded, and it does not bound how many copies of that expansion exist at once. Quote capacity as cap × expansion × concurrency, and make each of the three a number somebody chose.
- Why does decoding into a concrete struct cost so much less than decoding into an any?When the decoder is filling a typed struct it only materialises values that correspond to modelled fields; input it cannot place is scanned past. Decoding into `any` materialises whatever arrived — a `[]any` or `map[string]any` per level — so the sender chooses your allocation shape.
- The capacity planner asks what one request may cost. What number do you give?Not the byte cap. Give the product: byte cap times a measured expansion factor for the worst realistic payload, times the maximum requests in flight. Then name which of the three you control and propose a bound for the concurrency term, since it is usually the unbounded one.
- How do you confirm from a profile that decoding is the cost and not something downstream?Read a heap profile by allocation volume and see whether the decode path dominates the samples; take one during a burst for live bytes as well. If the heaviest sites sit after the decode, the payload is not the problem and a smaller body cap will not help.
- Does lowering the byte cap alone give a bounded peak?No. It shrinks one factor in a product of three. Without a bound on requests in flight, a smaller cap just means the poster needs more concurrent requests to reach the same peak — cheap for them, and invisible in your per-request numbers.
The byte cap weighs the flat-pack box. What fills the warehouse is the assembled furniture, and how much it expands to is printed on the sender's instructions, not on your scale.
saying these in an interview costs you the question
- Claims a 1 MB body cap bounds memory to 1 MB per request
- Ignores the number of requests in flight when quoting capacity
- Decodes untrusted payloads into any and calls it flexible
- Assumes deeply nested input costs proportional to its byte length
- Reaches for a CPU profile to explain an allocation spike
- Treats the spike as a leak because the heap grew