Your batch job's context.Context is cancelled and its cleanup writes now fail instantly. How do you make cleanup still run?
answer
- The cleanup inherited the cancellation
- Keep the values, drop the cancellation
- Background would lose your trace identifiers
- Detached still needs a bound of its own
- Closing files was never the context's job
basics
~20 sDo not run cleanup on the cancelled context. Derive one with context.WithoutCancel, which keeps the values but is never cancelled, bound it yourself, and use that for the closing writes. Files still need your own defer Close.
solid answer
~50 sCancellation propagates to every derived context, so the moment the job is cancelled every `ExecContext` on that context fails immediately with `context.Canceled` - including the audit row and the mark-run-failed update you wanted to write on the way out. The fix is to detach cleanup: `context.WithoutCancel(ctx)` returns a context that carries the same values but has no deadline and is never cancelled, so pass that to the closing writes rather than `context.Background()`, which would drop your request and trace identifiers. Then bound it yourself, because a detached context that hangs is a job that never exits. Separately, remember that cancellation is not cleanup at all: nothing in the context machinery closes an `*os.File`, closes `sql.Rows`, rolls back a transaction, or deletes a half-written temp file. Those remain your own `defer f.Close()` and rollback paths, and their absence shows up as open descriptors that never fall after a cancel.
code
go · 17 linesfunc processFile(ctx context.Context, db *sql.DB, path string) error {
f, err := os.Open(path)
if err != nil {
return err
}
defer f.Close() // cancellation never closes this for you
err = transform(ctx, f)
// Keeps the parent's values, but is never cancelled; bound it yourself.
cleanupCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), 5*time.Second)
defer cancel()
if _, cerr := db.ExecContext(cleanupCtx, recordRun, path, errText(err)); cerr != nil {
log.Printf("record run %s: %v", path, cerr)
}
return err // the original error, not the cleanup's
}go deeper
Know that a cancelled context makes every later database or HTTP call on it fail at once, so anything you want to happen after cancellation cannot use that same context.
Explain what context.WithoutCancel returns - the parent's values, no deadline, never cancelled - and why it is preferable to context.Background for cleanup work that should still be traceable.
Show the whole shutdown ordering: stop taking work, let the defers release resources, record the outcome on a detached bounded context, and return the original error rather than the cleanup's. Name the descriptor-count check that proves it worked.
Own what the system promises on cancel: which finishing work is guaranteed, how long it may take, and what is allowed to be abandoned. Anything longer than a few seconds of tail work belongs in durable state, not smuggled past shutdown on a detached context.
## The failure A batch job walks a tree of input files, transforms each one, and records progress in a database. On the way out - success or failure - it writes a run record. It looks like this: the run context is passed everywhere, the writes use `ExecContext`, and the final record is written in a deferred function. An operator cancels the job. The final record is never written. Worse, the error the job reports is `context.Canceled` from the *cleanup* write, which masks whatever real error caused the failure. The job leaves no trace of why it stopped. The reason is simple and it is the leaf's whole point: `database/sql` honours contexts. Cancelling the parent cancels every derived context immediately, and `ExecContext` on a cancelled context does not even attempt the statement - it returns the context error before touching the pool. ## Detaching cleanup properly `context.WithoutCancel(parent)` returns a context that is **not** cancelled when the parent is, has no deadline, and still returns the parent's values from `Value`. That last part is why it beats `context.Background()`: your trace identifier, request identifier, tenant and logger all live in those values, and starting from `Background` silently drops them so the cleanup write lands with no correlation to the run that produced it. But a context that can never be cancelled is a new hazard. If the database is the reason the job is being cancelled in the first place, an unbounded cleanup write hangs and the process never exits - and now the operator has a job that ignores both the cancel and the second cancel. So put your own bound on the detached context and let the cleanup fail loudly if it cannot complete. The pattern is: detach, bound, do the small closing writes, log if they fail, return the original error. Keep the detached work small and finite. `WithoutCancel` is for a handful of finishing operations - write the run record, release a lock row, delete a temp object. It is not a way to smuggle a long tail of work past a shutdown; anything long enough to care about deserves its own queue and its own lifecycle. ## Cancellation is not cleanup The second half of the answer is the one candidates skip. A context cancellation closes a channel. It does not: - close an `*os.File` - the descriptor stays open until your `defer f.Close()` runs, which means it stays open forever if the goroutine is parked in a read; - close `sql.Rows` - unclosed rows hold a connection out of the pool; - roll back a transaction - although a cancelled `BeginTx` context does cause `database/sql` to roll back the transaction for you, which is one of the few places the machinery does act; - delete a partially written output file, or undo anything already committed. The diagnostic that reveals this is blunt and worth naming: after the cancel, watch the process's open descriptors. If `lsof` shows the count flat or climbing while the job claims to be shutting down, some goroutine is still parked with a file open, and no amount of context work will fix it - the deferred `Close` is unreachable because the read above it never returned. There is a runtime finalizer on `*os.File` that eventually closes an unreachable file, but it is a backstop against leaks in the abstract, not a resource strategy: it runs whenever the collector gets to it, which may be long after you have run out of descriptors. ## Ordering the shutdown A clean shutdown path in this shape of program is layered, and saying the layers out loud is what separates a senior answer: 1. **Notice** the cancellation at the loop level, so no new unit of work starts. 2. **Release resources** through the defers already in place, which requires that every blocking call above them is bounded (chunked reads, deadlines where the descriptor allows, timeouts on outbound calls). 3. **Record the outcome** on a detached, bounded context, so the audit survives the very cancellation it is describing. 4. **Return the original error**, not the cleanup's `context.Canceled`, so the reason for stopping is not overwritten by a symptom of stopping. Step 4 is a small thing that shows up constantly in incident reviews: the cleanup error shadows the cause, and the postmortem starts from the wrong end.
- Why not just use context.Background() for the cleanup writes?It works for cancellation, but it throws away the parent's values - trace and request identifiers, tenant, the logger stored in the context. The cleanup record then lands with no correlation to the run it describes. `context.WithoutCancel` keeps the values and drops only the cancellation, which is exactly the split you want.
- What still needs to be closed by hand after a cancellation?Everything you opened: `*os.File` handles, `sql.Rows`, response bodies, temp files and directories. Cancelling a context closes a channel and nothing else. The one thing `database/sql` does for you is roll back a transaction whose BeginTx context is cancelled.
- How would you tell from outside that a cancelled job is not releasing resources?Watch the process's open descriptor count with `lsof` after the cancel. In a clean shutdown it falls as the defers run. If it stays flat, goroutines are still parked in blocking calls above their `defer Close`, and the goroutine profile will show exactly where.
saying these in an interview costs you the question
- Runs cleanup writes on the same cancelled context
- Reaches for context.Background and loses request-scoped values
- Leaves the detached cleanup context unbounded
- Says cancellation closes files and rows for you
- Returns the cleanup's context.Canceled instead of the original error