A Go CLI fans out to six internal APIs in one run. Where should each call's deadline live, and who owns it?
answer
- a ceiling and a budget are different jobs
- per call, never per invocation
- one client per dependency
- the slowest dependency should not set everyone's limit
- on-call needs a screw to turn
basics
~20 sUse both placements for different jobs: one http.Client per dependency whose Timeout is a ceiling owned by whoever holds that dependency's latency budget, plus one absolute deadline for the whole invocation carried in a context and passed to every request.
solid answer
~60 sGo gives you two places to put a deadline and they answer different questions. `http.Client.Timeout` is a per-call ceiling baked into the code that owns a dependency — it cannot express "the whole run finishes in five seconds", because six sequential calls at five seconds each is thirty. A `context.WithDeadline` created once in `main` and passed to every `http.NewRequestWithContext` *is* the invocation budget, and it is owned by the caller. I use both: a client per dependency with a ceiling derived from that dependency's measured p99 plus headroom, and one absolute deadline for the run that no call can outlive. What I refuse is a single shared client with one timeout for all six, which forces the slowest dependency's budget onto the fastest and destroys your ability to fail fast on the one that should always be quick. Ownership follows the same split: the team holding a dependency's latency budget sets its ceiling, and on-call can force it down after an incident — which is the argument for building all six clients in one factory fed by config, rather than six literals scattered across the code.
code
go · 10 lines// Ceilings owned by whoever holds each dependency's latency budget.
var billing = &http.Client{Timeout: 2 * time.Second}
var search = &http.Client{Timeout: 800 * time.Millisecond}
func run(parent context.Context) error {
// One budget for the whole command, owned by the caller.
ctx, cancel := context.WithTimeout(parent, 3*time.Second)
defer cancel()
return fanOut(ctx, billing, search)
}go deeper
Know that there are two places a deadline can live — on the client you built, or on the request you send — and that six calls under a per-call limit can still add up to a long wait.
Explain why a client field cannot express a budget for a unit of work, and show the mechanics of one derived deadline threaded through every request in a fan-out.
Argue for a client per dependency with limits derived from measured latency, and describe what a too-tight budget does to error rates and to load on the dependency you are already struggling with.
Own the policy: who sets each ceiling, how on-call overrules it without a code review, whether expiry cancels sibling calls in a fan-out, and what you accept in exchange for making the values configurable.
## Two placements, two meanings The mechanical choice is small and the consequences are not. **`http.Client.Timeout`** is a field on a shared object. It applies to every call that client makes, individually. It is invisible at the call site, it cannot be raised by a caller, and it is a compile-time or start-up decision. **A deadline in a `context.Context`**, attached with `http.NewRequestWithContext`, is per call and caller-controlled. It can be tightened per call, derived from a parent so that one deadline governs a whole tree of calls, and cancelled explicitly. Both are enforced; the earlier one wins. So they are not alternatives so much as a ceiling and a budget. ## The fact that decides the design `Client.Timeout` is **per call, not per invocation**. A CLI that hits six APIs sequentially with one client at five seconds each has a worst case near thirty seconds and never trips a timeout. If the product requirement is "this command answers in five seconds or tells the user it could not", no client field expresses it. Only one absolute deadline, derived once and passed down, does. This is the point most candidates miss, and it is the reason "we set timeouts everywhere" and "the command sometimes takes half a minute" are both true in the same codebase. Conversely a context deadline alone is not enough. A careless caller can pass a two-minute budget to a dependency that has never legitimately taken more than 80 milliseconds, and now a stuck dependency holds a connection and a goroutine for two minutes because nobody told it not to. The client's ceiling is the guard rail the dependency's owner puts up against callers they cannot review. ## The shape I argue for - **One `http.Client` per dependency**, built at start-up in one place, each with a `Timeout` set from that dependency's measured p99 plus headroom. Reusing one client per dependency also keeps connection reuse per host where it belongs. - **One absolute deadline per invocation**, created in `main` (or per request, in a server) and threaded into every `http.NewRequestWithContext` call. - **Per-call tightening** where a specific call deserves less than its dependency's ceiling — a health probe, a best-effort enrichment. - **No package-level client, no shared client across dependencies.** A single client for all six couples them: the slowest dependency dictates the ceiling, and a fast one that hangs now hangs for the slow one's budget. ## The fan-out judgment call When the six calls run concurrently and share one deadline, expiry kills all of them. Whether that is right is a product decision, not a technical one. If the command's output is all-or-nothing, one shared deadline is exactly correct — cancelling the siblings the moment the answer cannot be complete saves everyone work. If partial output is useful ("five of six sections rendered, one timed out"), then each call needs a derived deadline of its own so one slow dependency degrades one section instead of failing the command. Decide it explicitly and write it down; the default that emerges from whichever context you happened to pass is not a decision. ## Numbers, and who gets to change them A timeout is a promise about how long you will wait, and setting it too low is worse than setting it too high: cap at 500 ms a dependency whose p99 is 900 ms and you have converted a slow dependency into a hard failure for a tenth of traffic, and if anything retries you have multiplied load on a system that was already struggling. Derive from measurement, add headroom, and alert on the *rate* of timeouts rather than on individual ones. Ownership then falls out naturally. The engineer who holds a dependency's latency budget sets its ceiling, because they are the one who knows its p99 and who will be asked why calls fail. The on-call SRE gets to overrule it during and after an incident — "nothing waits more than a second on search until it is fixed" — and that veto is only usable if the value lives somewhere they can change without a code review at 3am. That is the real argument for building the clients in one factory over a config struct: not elegance, but the ability to turn one screw under pressure. The cost is honest and worth stating: a config-driven timeout is one more untested value that can drift, so give it a sane default in code, log the effective values at start-up, and cover the factory with a test. ## What I would say in a review Ask three questions of any client code: what bounds a single call, what bounds the whole unit of work, and who can change either without a deploy. Code that answers all three is fine whatever the numbers are. Code that answers none of them will eventually spend an unbounded amount of time waiting for somebody else's outage.
- Why can't one http.Client.Timeout express "the whole run finishes in five seconds"?Because it is applied to each call independently. Six sequential calls under a five-second client limit have a worst case near thirty seconds and never trip it. Only an absolute deadline derived once and passed to every request bounds the unit of work as a whole.
- On-call wants every outbound call capped at one second during an incident. Where do you put that?In the one place all clients are constructed, fed by config, so a single value changes all six without editing call sites. Keep a sane default in code, log the effective values at start-up, and test the factory — a knob nobody can find at 3am is not a control.
- When is a timeout that is too short worse than no timeout?When it sits under the dependency's real p99. Capping at 500 ms something that legitimately takes 900 ms fails a tenth of traffic outright, and if anything retries you amplify load on a system already in trouble. Derive from measurement plus headroom and alert on the timeout rate.
- In a concurrent fan-out, should one expiry cancel the sibling calls?It depends on whether partial output is useful. All-or-nothing commands should share one deadline so expiry abandons every sibling immediately. If the command can render five of six sections, give each call its own derived deadline so one slow dependency degrades one section rather than the whole run.
saying these in an interview costs you the question
- Uses one shared client and one timeout for every dependency
- Thinks a per-call client limit bounds the whole invocation
- Picks timeout values by intuition rather than measured latency
- Hardcodes limits nobody can change during an incident
- Sets a budget below the dependency's own p99