skip to content

A Go service starts cleanly but every request fails instantly with a deadline error; the timeout comes from an env var. How do you diagnose and fix it?

level: seniorimportance: should knowfreq 48%

answer

  1. instant failure means no time was left
  2. what is the zero value of a duration
  3. who threw away the parse error
  4. a bare number is not a duration
  5. zero deadline equals a deadline of now

basics

~20 s

The timeout almost certainly resolved to a zero time.Duration, because the variable was absent or malformed and the parse error was discarded. context.WithTimeout with zero builds an already-expired deadline. Fix it by checking the parse error and validating the value before the server starts.

solid answer

~50 s

This is the signature of a zero `time.Duration`. If the variable is missing, or holds something `time.ParseDuration` rejects such as `30` with no unit, and the error is dropped with `_`, the setting silently becomes `0`. `context.WithTimeout(ctx, 0)` then computes a deadline of now, so the context is already expired and every handler fails before doing any work. I would confirm it by logging the resolved configuration at startup — the value prints as `0s` and the case closes — then reproduce it in a table-driven test over the precedence cases. The fix is structural, not a bigger default: check the error from every `time.ParseDuration` and `strconv` call, range-check the result, collect the failures with `errors.Join`, and exit nonzero from `main` before anything listens. Note that zero is not universally 'no timeout' — `http.Client.Timeout` of zero means unlimited, which is why zero can never be a reliable sentinel.

code

go · 9 lines
go
// wrong: an absent or malformed value silently becomes 0
var timeout, _ = time.ParseDuration(os.Getenv("REQUEST_TIMEOUT"))

func handle(w http.ResponseWriter, r *http.Request) {
	ctx, cancel := context.WithTimeout(r.Context(), timeout)
	defer cancel()
	// timeout == 0: the deadline is now, so ctx is already expired here
	_ = ctx
}

go deeper

for a junior

Know that a failed time.ParseDuration returns zero along with the error, and that ignoring the error leaves you holding that zero. Always check the error when turning a string into a typed value.

for a middle

Explain the chain precisely: zero duration, deadline computed as now, context already cancelled with DeadlineExceeded before any work runs. Know that time.ParseDuration rejects a value with no unit.

for a senior

Demonstrate the diagnosis path and the structural fix: log the effective config at startup, reproduce with a table-driven test over absent and malformed values, then validate every setting and refuse to start on a bad one.

for a principal

Decide the policy this bug argues for: which settings are allowed a default and which must abort startup, and how you keep a fail-fast service from turning one typo into a fleet-wide crash loop during rollout.

## Reading the symptom "Starts cleanly, then fails instantly on every request" is a very specific shape. Instant and universal rules out load, dependencies and network — nothing had time to happen. A deadline error with no elapsed time means the deadline was already in the past when the request began. That points at one number: a duration of zero. ## Why zero produces this exact failure `context.WithTimeout(parent, d)` is defined as `context.WithDeadline(parent, time.Now().Add(d))`. With `d == 0` the deadline is *now*; the remaining time is not positive, so the returned context is cancelled immediately with `context.DeadlineExceeded`. Any code that does the honest thing — check `ctx.Err()`, or pass the context to a call that respects it — fails before doing work. Nothing panics, nothing logs an error at startup, and the service looks healthy from the outside until the first request. ## How the zero got there Two routes, both extremely common. **The discarded error.** `time.ParseDuration` returns `(0, err)` on failure, and `d, _ := time.ParseDuration(os.Getenv("REQUEST_TIMEOUT"))` throws the error away and keeps the zero. The value that triggers it is usually not exotic: `time.ParseDuration` **requires a unit**, so `30` is invalid where `30s` is fine. Only `0` may be written without one. An operator who typed a plain number, meaning seconds, has produced exactly this outage. `strconv.Atoi` behaves the same way — `0` plus an error — so an ignored error there yields a zero threshold or a zero worker count. **The absent variable.** The variable was never set in the new environment, the code took the zero value of the field, and nothing objected. This is the harder version, because there is no typo to find and the deployment looks correct. It is also why the zero value of a numeric setting is dangerous: in Go a struct field that was never populated and a field deliberately set to zero are indistinguishable at the point of use. ## Diagnosis, in order 1. **Print the resolved configuration at startup**, secrets redacted. A line reading `request_timeout=0s` ends the investigation in seconds. If the service does not do this yet, that is the first fix regardless of this bug. 2. **Check the environment of the running process**, not the deployment template — the two disagree more often than anyone expects, particularly when a variable name was changed. 3. **Reproduce with a table-driven test** over the resolution function: a row for the value absent, a row for a malformed value like `30`, a row for a valid `250ms`, each asserting the resolved struct. The absent and malformed rows should assert an error, and before the fix they will fail by returning zero. 4. **Grep for discarded errors** around parsing. `, _ :=` next to `ParseDuration`, `Atoi`, `ParseFloat` or `ParseBool` is the pattern; `go vet` will not flag it, because ignoring a returned error is legal Go. ## The fix Raising the default is not a fix; it hides the next occurrence. Three changes make the class of bug impossible: **Check every conversion.** Every string coming from the environment or the command line becomes a typed value through exactly one call whose error is checked and wrapped with the setting's name and the offending text: `fmt.Errorf("%s=%q: %w", name, raw, err)`. The name and value in the message are the difference between a five-minute fix and an hour. **Validate ranges, not just syntax.** A syntactically valid `0s` is still not a legal request timeout, and a rollout percentage of `140` parses perfectly. Assert what the setting means: positive durations, percentages within 0 to 100, an address that is non-empty. **Report everything at once and refuse to start.** Collect the failures into a slice and return `errors.Join(errs...)`, which is nil when the slice is empty. The engineer bringing the service up in a new environment then sees all four broken settings on the first attempt instead of discovering them one deploy at a time. Do the validation in `main`, after every source has been resolved and **before** the process listens or dials anything, then print the error to stderr and exit nonzero. Keep the exit in `main`: `os.Exit` does not run deferred functions, so a helper that exits skips every cleanup its callers registered. ## Why zero cannot be a sentinel One last trap worth stating explicitly, because it defeats the instinct to "just treat zero as unlimited": zero means opposite things in the standard library. `http.Client.Timeout` of zero means **no timeout at all**, and `http.Server.ReadTimeout` of zero likewise means unlimited — while `context.WithTimeout` with zero means **already expired**. A single config value of `0` fed to both produces a request that never times out at the client and always times out at the handler. Validate the value; never infer intent from zero.

  • Why does time.ParseDuration reject "30"?
    It requires a unit on every non-zero value: `30s`, `1500ms`, `1h30m` are valid, plain `30` is not, and only `0` may be written unitless. The rule exists because a bare number is ambiguous — Go's own `time.Duration` counts nanoseconds, while a human typing `30` almost always means seconds. Rejecting it prevents a silent factor-of-a-billion error.
  • Where exactly in main should configuration validation run?
    After every source has been resolved and before the process does anything observable — before it listens, dials a database, or registers with service discovery. A bad configuration should be a non-start with a nonzero exit status, not a process that accepts traffic and fails each request. Print the error to stderr naming the setting and the offending value.
  • Why keep os.Exit in main rather than calling it from a validation helper?
    `os.Exit` terminates the process immediately and does not run deferred functions, so any cleanup a caller registered with `defer` is skipped and buffered output can be lost. Helpers should return errors and let `main` decide to exit, which also keeps them testable — a helper that exits cannot be unit-tested.
  • Would the race detector or go vet have caught this?
    No. There is no data race, and discarding a returned error with `_` is legal Go that `go vet` does not report. This class of bug is caught by tests over the resolution function and by validating values at startup — not by the toolchain's dynamic or static checks.

saying these in an interview costs you the question

  • Discards the error from time.ParseDuration with a blank identifier
  • Assumes a zero duration always means 'no timeout'
  • Fixes it by raising the compiled-in default value
  • Validates configuration lazily on the first request
  • Boots with a guessed default for a required setting and logs a warning