How do you decide whether a shared Go client package sets its own context.WithTimeout or only honours the caller's deadline?
answer
- a library can only shorten, never extend
- ask ctx.Deadline() and check ok
- the constant is inherited by every importer
- configurable field with a documented default
- caller owns overall, library owns per-attempt
basics
~10 sTreat the caller's deadline as authoritative and any package-set timeout as a configurable cap that mainly applies when the caller supplied none. Read ctx.Deadline(); a package can only shorten a budget, never extend one.
solid answer
~50 sStart from the mechanism: `context.WithTimeout` inside a library can only shorten what the caller gave you, so a constant written there is a **cap**, never a guarantee of time. That makes it a question of ownership. The caller owns its latency budget, so its deadline is authoritative and is passed through unchanged whenever it exists. The library owns protection against a caller who supplied none - `ctx.Deadline()` returning `ok == false` is the case a default cap exists for. I make that cap a documented field on the client struct rather than a hard-coded constant, because a number compiled into a package forty services import is a latency policy all forty teams inherit without review. If the package retries, the caller's context is the overall budget and each attempt gets a configurable slice of it, never a fresh unbounded one.
go deeper
The takeaway to hold onto is that a timeout written inside a library is an upper bound, not a reservation. If your caller had less time, you get less time.
Be able to describe the mechanics behind the decision: reading ctx.Deadline() and its ok flag, deriving a shorter child rather than a fresh root, and returning ctx.Err() so callers can tell an expiry from a failure.
Demonstrate that you would find the binding layer from evidence - per-hop timings and pool gauges - before changing any number, and that you know a timed-out caller still leaves your library holding resources until your own cap fires.
Own the position: whose SLO the number encodes, why it is configuration rather than a constant, how a dependency bump that changes it gets reviewed, and when guarding a shared resource justifies a cap importers cannot raise.
## What the mechanism decides for you Before any judgment, one fact narrows the choice: a derived context can only make a deadline **earlier**. `context.WithTimeout(ctx, 30*time.Second)` inside a package, called with a context that has 500 ms left, produces a context with 500 ms left. So the two options are not symmetric. "Set my own timeout" cannot mean "guarantee myself 30 seconds"; it can only ever mean "refuse to run longer than 30 seconds". A package-set timeout is a ceiling, and the only situation where the ceiling is the *operative* number is when the caller's budget is larger - or absent. That reframes the decision as: **who owns the ceiling, and when does the ceiling matter?** ## The three cases, and what to do in each **The caller supplied a deadline tighter than your cap.** Nothing to decide - the caller wins, automatically. Pass the context through. The only thing to get right is not *degrading* it: do not start work on `context.Background()` internally, and do not do the expensive part before you derive anything. **The caller supplied no deadline at all** - `ctx.Deadline()` returns `ok == false`. This is the case a library default genuinely exists for. Without one, a hung dependency means a goroutine, a connection and a pool slot held forever, and the failure mode is a slow resource leak rather than an error. A default cap here is defensible and I would ship one. **The caller's deadline is longer than anything your dependency can usefully survive.** This is the contested case. Imposing your cap overrules a caller who may have deliberately chosen a long budget for a batch job. My default is to allow it but make it visible: a field on the client, documented, with the default value stated in the package docs, so the batch team can raise it for their client instance without forking your package. ## Why the constant is a policy question A number compiled into a package that forty services import is a latency and capacity policy that forty teams inherit without ever having reviewed it. It shows up in three ways they will notice: - **Capacity.** If your cap outlives the front door's budget, every timed-out request still holds a connection and a pool slot for the remainder of your cap. The user gave up; your library did not. Under load that inflates effective service time and saturates pools - the classic shape of a cascading timeout incident, where the postmortem's root cause is one constant in a package nobody on the incident call maintains. - **Attribution.** When your cap fires, the error surfaces from your package. Every importing team reads it as "the client library timed out" and files it against you, whether or not your cap was the binding one. - **Change control.** Raising or lowering that constant is a behaviour change shipped through a dependency bump. Someone must own that upgrade and its blast radius - which is precisely the sort of decision an SRE reviewing the incident is entitled to overrule you on. The design that survives review: **a configurable field with a documented default, plus per-hop timings you export**, so that when a deadline fires the histogram shows which layer's budget was binding instead of leaving the postmortem to guess. ## Retries: per-attempt versus overall If the package retries, it is no longer just honouring a budget - it is *spending* one, and the split becomes explicit. The rule I hold to is that **the caller's context is the overall budget** and it is never restarted, extended, or replaced between attempts. Each attempt derives a shorter child from it. Two consequences follow directly from the mechanism: - The loop must stop when the parent is done, and the check is on the parent, not on the per-attempt child, because a per-attempt expiry and the overall one look identical at the error value. Retrying after the caller's budget is gone is pure waste that the caller cannot see. - The per-attempt number is not derivable by the library alone. It is a function of the caller's tolerance for tail latency versus its tolerance for a single slow dependency, so it belongs in configuration next to the cap, not in a constant. A package that retries silently, without either of those, is also making a load decision on its importers' behalf: N attempts multiplies the traffic every importing service sends to a struggling dependency. ## What I actually ship A client struct with a `Timeout` field defaulting to something conservative, documented as applying only when the caller's context has no earlier deadline; the caller's context passed to every network call unchanged; a per-attempt budget and attempt cap that are configurable and default to something modest; `ctx.Err()` returned rather than wrapped into an opaque error; and a metric per attempt so the caller can see which budget bound the call. The judgment I am explicitly *not* making for my importers is what their latency SLO is. ## The position to be able to defend either way The opposite call is legitimate in one situation: a package guarding a shared resource whose availability is everyone's problem - a connection to a database with a fixed pool, say. There, a hard cap that importers cannot raise is a protection mechanism rather than policy overreach, and "no caller may hold this for more than N seconds" is the whole point. State which of the two you are building, because the review comment you will get is exactly that question.
- What specifically do you do when ctx.Deadline() reports ok == false inside your package?Apply the configured cap - that is the case it exists for. An absent deadline means no one upstream bounded this call, so a hung dependency would hold a goroutine, a connection and a pool slot indefinitely. I also treat it as a signal worth counting: a rising share of budget-less calls usually means a caller lost its plumbing somewhere.
- How do you split an overall budget across retry attempts inside such a package?The caller's context stays the overall budget and is never reset; each attempt derives a shorter child from it. The loop checks the parent's `Done()` before every attempt, so it stops as soon as the caller's time is gone rather than after a fixed count. Both the per-attempt slice and the attempt cap are configuration, because they trade tail latency against dependency slowness in a way only the caller can weigh.
- An SRE says your library's 30-second cap caused a cascading incident behind a 5-second front door. What is your response?They are right that the cap was binding on the wrong side. The fix is not a smaller constant but a structural one: the front door's deadline must actually reach my package - if requests arrived with 5 s left, my 30 s could never have applied - so the first question is where the context chain was broken. Then I make the cap configurable and export per-attempt timings, so the next histogram shows which layer bound the call.
- When is a hard cap that importers cannot raise the right call?When the package guards a shared resource whose exhaustion is everyone's problem - a fixed connection pool, a rate-limited upstream. There "no caller may hold this longer than N" is the protection the package exists to provide, and making it configurable would let one importer degrade every other. Say explicitly which of the two kinds of package you are building.
saying these in an interview costs you the question
- Thinks a library timeout guarantees itself that much time
- Hard-codes a constant and calls it an implementation detail
- Starts work on context.Background to protect its own operation
- Restarts or extends the caller's budget between retry attempts
- Swallows ctx.Err() into an opaque package-specific error