During a drain, how do you make a second SIGTERM stop a Go worker immediately instead of being swallowed?
answer
- the handler outlives the first signal
- registration disables the default kill
- an ignored operator is the real symptom
- unregister, or listen for the next one
- three endings, three distinct log lines
basics
~20 sWhile Go's signal handler stays registered, a second SIGTERM no longer kills the process, so it does nothing during the drain. Either unregister the handler when the drain starts, restoring the default kill, or watch for it and abort.
solid answer
~50 sRegistering for a signal disables its default disposition, so once shutdown has begun a second SIGTERM is delivered to the runtime and quietly dropped — the root context is already cancelled and nobody is listening. The operator presses the button again and nothing happens. Two ways to fix it. The blunt one: call the stop function that `signal.NotifyContext` returned as soon as the drain begins; that unregisters the handler and restores the default behaviour, so the next signal terminates the process the way the operating system would have. The controlled one: keep your own notification and select on it alongside the drain budget, so a second signal cancels the drain context, lets you log how many items you are abandoning, and exits with a distinct status. The blunt version is instant but silent; the controlled one tells you what was lost, but only runs if your own code is not wedged too.
code
go · 9 linessigCtx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
<-sigCtx.Done()
log.Print("signal received, draining")
stop() // another SIGTERM or ^C now terminates the process immediately
drainCtx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()go deeper
Know that once a Go program asks to be notified about a termination signal, that signal stops killing it by default. The first one starts your shutdown; a later one does nothing unless the program arranges for it.
Explain both remedies and what each costs: unregistering the handler restores the default kill instantly but silently, while watching for a second signal yourself gives you a log line and a distinct exit status but only works if your own code is still able to run.
Argue from the operator's seat. Say why a drain that ignores a repeated stop is an incident-time hazard, keep operator-abort and deadline-expiry as separate reported outcomes, and note that a wedged process may only be stoppable from outside.
Set the contract the whole fleet follows — a repeated stop signal must stop the process without waiting for the drain — and decide whether the standard is silent unregistration or a reported abort, so on-call behaviour is the same across services.
## Why the second signal does nothing Asking the Go runtime to deliver a signal to your program also turns off what that signal would otherwise do. That is what makes graceful shutdown possible in the first place: a termination signal stops killing the process and starts arriving as an event instead. The consequence people miss is that the disabling does not end when the first signal has been handled. Registration lasts until it is explicitly undone. So during the drain: 1. the first signal arrives, the root context is cancelled, the drain starts; 2. the operator, or the platform, sends a second signal; 3. the runtime delivers it into a channel that is buffered by one and that nobody is reading any more; 4. the process keeps draining as if nothing happened. From the outside this is the worst moment for the program to stop responding. Someone is pressing the key again precisely because the drain is taking longer than they expected, and the program is choosing to ignore them. ## Option one: restore the default disposition `signal.NotifyContext` hands back a stop function alongside the context. Calling it unregisters the handler, and once no handler is registered, the signal reverts to its default effect — terminating the process. So the whole fix is to call it at the top of the shutdown path rather than only on the way out: ```go sigCtx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM) defer stop() <-sigCtx.Done() stop() // from here on, another SIGTERM or ^C kills the process outright // ... bounded drain ... ``` Calling the stop function twice is harmless, so keeping the `defer` is fine. The properties of this option: it is instant, it needs no extra state, and it works even if the drain itself is wedged in a way that your own code could not escape — the kill comes from outside your program. The cost is that it is completely silent. No log line, no count of what was in flight, no flush. ## Option two: handle it yourself If you want the second signal to be *reported* rather than merely obeyed, keep listening and treat it as another way for the drain to end: ```go select { case <-workersDone: log.Print("drain complete") case <-drainCtx.Done(): log.Printf("drain deadline: abandoning %d items", inflight.Load()) case <-abort: // a second signal arrived log.Printf("drain aborted by operator: abandoning %d items", inflight.Load()) } ``` Now the three ways a drain can end are three distinct branches, each with its own log line and its own exit status. On call, that difference is worth a lot: "we hit the deadline" and "a human cut it short" are different stories about the same incident. The risk is that this path is still your code. If the drain is blocked in a way that also blocks your abort handling — a goroutine wedged holding something the exit path needs — the operator's second signal is once again ignored, and the only thing left is the kill they cannot send because you are still registered for it. ## The pragmatic combination Many production programs do both, in order: on the second signal, log the abandoned count and unregister the handler; on the third, the operating system takes over. Whatever the arrangement, the requirement is the one an operator would state: **after the second signal, the process must be stoppable without waiting for the drain.** ## What this is not It is not a substitute for the drain budget. The budget bounds the ordinary case, where nobody is watching, and it is what makes automated deploys predictable. Second-signal handling covers the case where a human has decided the ordinary case is taking too long. A program needs both, and they should end up in different branches with different reported statuses so the two situations stay distinguishable afterwards. ## An uncatchable reminder There is always a last resort that no program can intercept, and the platform will use it when its grace period expires. Designing the second-signal path is about making the *human's* escape hatch work at the speed a human expects; it never removes the possibility that the process is stopped from outside without any of your code running.
- Why is the second signal dropped rather than queued up for later?The runtime delivers signals into a small buffered channel and never blocks to do it. During the drain nothing is reading that channel any more, so once the buffer holds one value further deliveries are discarded. Signals are notifications of "this happened", not a queue you can replay.
- Is it a problem to call the stop function returned by signal.NotifyContext twice?No. It is safe to call more than once, so calling it explicitly when the drain starts and keeping the deferred call for the normal exit path is fine, and is the usual way this is written.
- Which approach would you default to in a worker nobody watches interactively?Unregistering at the start of the drain. It costs one line, cannot itself get stuck, and the batch case is already covered by the budget. Handling the second signal in code is worth the extra machinery mainly where an operator regularly stops the process by hand and you want the abandoned count on record.
It is a fire alarm with the sounder disconnected while the drill runs: the second person to hit the button is not wrong, they simply cannot be heard until someone reconnects it.
saying these in an interview costs you the question
- Assumes a second SIGTERM always kills a Go process
- Thinks the handler is torn down after the first signal
- Expects extra signals to queue up and be handled later
- Offers no way out of a drain except waiting for the deadline
- Reports an operator abort and a deadline with the same status