How do you choose the maximum inbound message size for a Go service when a caller team's traffic already exceeds your candidate cap?
answer
- the cap is not the number that matters
- multiply by messages in flight
- measure the tail, not the mean
- log what it would have rejected
- move the outlier, not the ceiling
basics
~20 sDerive the cap from worst-case memory: the per-message limit times the messages in flight must fit the process budget. Measure real payload sizes, shadow-log what a candidate cap would reject, and move outliers to streaming rather than raising the ceiling.
solid answer
~50 sStart from arithmetic rather than a round number: worst-case memory is roughly the per-message cap multiplied by the messages that can be in flight at once, and that product has to fit inside the container's memory with room for everything else the process does. `GOMEMLIMIT` is a useful backstop, but it degrades into collector thrash rather than a clean rejection. Then get data: measure the real size distribution and shadow-log which requests a candidate cap would refuse, so you learn who breaks without breaking them. When one caller genuinely needs more, the answer is usually chunking, streaming, or passing a reference to the payload rather than a higher global ceiling. Publish the limit with an actionable error, and remember that raising it spends fleet-wide memory to serve one caller: that is the tradeoff you own, and the one you can be overruled on.
go deeper
The takeaway is that a size limit exists and is a deliberate choice, not a default. Be able to say why unbounded input is unsafe even if you are not the one picking the number.
Be ready to do the arithmetic: per-message cap times messages in flight against the process memory budget, and explain why lowering concurrency is an alternative to lowering the cap.
Show the rollout: measure the size distribution, shadow-log what a candidate cap would refuse, then enforce in steps with rejections monitored per caller.
Own the tradeoff explicitly, including who you are prepared to refuse and why, the migration path you offer them, and the conditions under which you would be right to be overruled.
## Why this is a judgment call and not a lookup A maximum inbound message size looks like a constant and behaves like a policy. Set it too high and one message can consume memory the whole process needs; set it too low and you refuse work someone legitimately sends. Nobody can hand you the right number, because it depends on your memory budget, your concurrency, your callers, and how much you are willing to break to be safe. ## Step 1: the arithmetic The number that matters is not the cap, it is the product: ``` worst-case bytes ~= per-message cap x messages in flight ``` For a server that runs a goroutine per connection and buffers a whole message, "messages in flight" is the connection limit. A 4 MiB cap sounds modest until it meets ten thousand connections, at which point it is 40 GB. Two consequences follow. First, the per-message cap and the concurrency limit are one decision, not two — if you cannot lower the cap, lower the number of simultaneous readers, for example with an admission semaphore around the buffering step. Second, the budget must leave headroom for everything else: request handling, caches, the collector's own working space. `GOMEMLIMIT` belongs in this conversation as a backstop rather than a solution. It is a soft limit: as the heap approaches it the collector runs harder and more often. That converts an out-of-memory kill into a service that is alive but slow, which is usually preferable but is still an outage. The cap is what prevents you needing it. ## Step 2: measure before you decide Guessing the distribution is how caps get set to the wrong power of two. Instrument the sizes you actually receive, keep a histogram, and look at the tail, not the mean. Then run the cheapest experiment available: **enforce nothing, log everything a candidate cap would have rejected**, with the caller identity attached, for long enough to cover weekly and monthly peaks. That report is the whole political content of the decision — it turns "we think this will be fine" into a list of exactly which teams to talk to. Roll out in the same spirit: a generous limit first, then tighten in steps, each step preceded by the same shadow report. ## Step 3: what to do about the caller who does not fit The pressure will always be to raise the ceiling for everyone because one caller sends big payloads. Resist it, because that spends fleet-wide memory to serve one workload, and the ceiling never comes back down. Better answers, roughly in order: - **Chunk or stream.** If the payload can be processed incrementally, it does not need to be buffered, and the cap stops being relevant to it at all. This is often a smaller change than it sounds. - **Pass a reference.** Large blobs go to storage; the message carries a pointer. This also usually improves retries and observability. - **A separate endpoint or deployment** with its own limits and its own memory budget, so the risk is contained to the workload that needs it. - **A per-caller limit**, if your framing supports identifying the caller before the body arrives. Powerful, but it adds configuration that has to be owned and reviewed, and a per-caller exception has a way of becoming permanent. Raising the global cap is the last option, and if you take it, you take it with the arithmetic from step 1 redone and, usually, the concurrency limit lowered to pay for it. ## Step 4: make the limit part of the contract A limit nobody can discover is an outage waiting to be debugged by someone else. Document the number, return an error that names it and says what to do instead, and version it: tightening a published limit is a breaking change and deserves notice, a deprecation window, and the shadow report as evidence. On the operational side, count rejections as a first-class metric with the caller attached. A rising rejection rate is either an attack or a caller that has outgrown the contract, and you want to know which within minutes. Rejections should also be cheap: reject on the header, before allocating, and close the connection rather than draining it. ## Step 5: own the tradeoff out loud The honest framing to give a room is: "This cap rejects some legitimate large messages in exchange for a bound on the memory any one peer can make us spend. Here is the worst-case memory at the current number, here is who it would have refused last month, and here is the migration path for them." That is a decision the service owner makes because the service owner carries the pager for the memory — and it is a decision that can be legitimately overruled if the business value of the refused traffic exceeds the availability risk. What is not defensible is having no number, or having one nobody can explain.
- Does GOMEMLIMIT remove the need for a message size cap?No. It is a soft limit: as the heap approaches it the collector runs harder rather than refusing the allocation, so an oversized message turns into collector thrash and latency instead of an out-of-memory kill. That is a better failure but still a failure. The cap prevents the pressure; GOMEMLIMIT only shapes what happens if prevention fails.
- How do you tighten an already-published limit without breaking callers?Treat it as a breaking contract change. Run the new number in shadow first, logging what it would refuse with the caller attached; share that list with the affected teams; give a deprecation window with a migration path such as chunking or a reference to stored data; then enforce. The evidence is what makes the conversation short.
- When is a per-caller limit worth the complexity?When one workload is genuinely different in kind and you can identify the caller before the body arrives. It contains the memory risk to that caller instead of the fleet. The cost is configuration that must be owned, reviewed and expired, since per-caller exceptions tend to become permanent and invisible.
- What do you monitor once the cap is live?The rejection rate broken down by caller, the observed size distribution against the cap so you can see the tail approaching it, and the worst-case memory figure recomputed whenever the concurrency limit changes. A sudden rejection spike is either an attack or a caller that outgrew the contract, and the label tells you which.
saying these in an interview costs you the question
- Picks a round number with no memory arithmetic behind it
- Ignores concurrency, so the per-message cap is the only bound quoted
- Raises the global limit to accommodate one caller
- Enforces a new limit in production without a shadow measurement first
- Treats GOMEMLIMIT as a substitute for a cap
- Leaves the limit undocumented and the rejection error unactionable