A certificate-renewal loop that ticks every minute now takes longer than a minute per run. How do you make that visible and bound it?
answer
- the loop looks healthy either way
- measure gaps, not just errors
- bound the run, not the loop
- defer in a loop body is per function
basics
~20 sMeasure the gap between successive ticks and export a missed-run counter plus a run-duration timing, then bound each run with its own context.WithTimeout derived from the loop's context, and shrink the work per run instead of firing overlapping goroutines.
solid answer
~60 sFirst make it visible. On each tick, record the gap since the previous tick and add the surplus over the interval to a missed-run counter, and time every run — a loop that should tick 60 times an hour and ticks 20 says so immediately, where error counts and success logs say nothing. Then bound the run rather than the loop: derive `runCtx, cancel := context.WithTimeout(ctx, …)` per tick, pass it into the work so a hung certificate authority call cannot own the loop forever, and call `cancel` before the next iteration — a bare `defer cancel()` in a loop body holds every context until the enclosing function returns, which `go vet` flags. Then shrink the work: filter to certificates whose `NotAfter` is inside the renewal window instead of sweeping all of them, and parallelise inside a single run with a bounded number of goroutines. Widening the interval to something you can actually meet is a legitimate fix too. What I would not do is fire each tick into its own goroutine — that trades skipped runs for unbounded overlap.
code
go · 8 linescase <-ticker.C:
func() {
runCtx, cancel := context.WithTimeout(ctx, 45*time.Second)
defer cancel()
if err := renewExpiring(runCtx); err != nil {
slog.Error("renewal sweep failed", "err", err)
}
}()go deeper
Know that a run taking longer than the interval simply means fewer runs, and that a context with a deadline is how you stop one run from lasting forever.
Be able to write the per-run timeout correctly, including why cancel must be scoped to the iteration and why the deadline must be derived from the loop's own context.
Lead with detection: name the gap-versus-interval measurement, the alert you would set on it, and then argue the tradeoff between shrinking the run and widening the interval.
Decide what the platform promises. Whether periodic work is allowed to skip, what the missed-run budget is, and who is paged when it is exceeded are policy calls that belong in one place, not per daemon.
## Why nothing alerted A periodic loop whose runs outgrow its interval fails in the quietest possible way. Every run that happens still succeeds. No goroutine leaks, no panic, no error rate moves. The only thing that changed is the *rate*, and rate is the one thing daemons typically do not measure about themselves. The engineer whose nightly refresh started overlapping itself as the dataset grew usually discovers it from the downstream consequence — a certificate that expired, a cache that was stale for an hour — not from the loop. So the first move is instrumentation, and it is cheap. ## Make the schedule observable Three numbers are enough: 1. **The gap between successive ticks.** Record `time.Now()` at each receive; compare the gap with the configured interval; anything beyond roughly two intervals means at least one scheduled run never happened. Accumulate the surplus into a counter. 2. **The duration of each run.** A histogram or even a last-run gauge tells you whether the loop is close to its ceiling before it goes over. 3. **The outcome of each run**, so a run that fails fast is not mistaken for a healthy fast run. Alert on the first. "This loop has skipped runs in the last N minutes" is an actionable, low-noise signal that catches the growth-driven version of this problem months before a certificate expires. ## Bound the run, not the loop The second move is to stop a single bad run from owning the schedule indefinitely. Give each run its own deadline derived from the loop's context, so cancellation still propagates from shutdown *and* a hung dependency cannot pin the loop: ```go case <-ticker.C: func() { runCtx, cancel := context.WithTimeout(ctx, 45*time.Second) defer cancel() if err := renewExpiring(runCtx); err != nil { slog.Error("renewal sweep failed", "err", err) } }() ``` Two details matter here. The timeout context is derived from the loop's `ctx`, so cancelling the daemon cancels the in-flight run too — otherwise shutdown waits for a run that no longer needs to finish. And the `cancel` is scoped: the `defer` sits inside a function literal that returns each iteration. Writing `defer cancel()` directly in the loop body is a classic Go bug — `defer` is per function, not per iteration, so every iteration's cancel function piles up until the loop's enclosing function returns, which for a daemon is never. `go vet`'s `lostcancel` check catches the variant where `cancel` is never called at all; the piling-up variant it will not catch for you. The timeout must of course be honoured by the work. A `renewExpiring` that takes `context.Context` and ignores it — no `ctx` on the HTTP requests, no `ctx.Err()` check between certificates — gives you a deadline that expires while the run carries on regardless. ## Then make the run fit Instrumentation and deadlines tell the truth about the problem; they do not solve it. The run has outgrown the interval, and there are only three honest responses: - **Do less per run.** The nightly-refresh version of this problem is almost always a full sweep that should be a filtered one: select only the certificates whose `NotAfter` falls inside the renewal window, or those changed since the last run, rather than every certificate in the store. - **Do it faster.** Parallelise *inside* one run with a bounded worker count, so the loop still has exactly one run in flight but that run uses the machine. Bounded is the operative word: the downstream certificate authority has its own rate limit. - **Slow the loop down.** If a full sweep genuinely takes four minutes, a one-minute interval is a configuration that has never been true. Setting it to five minutes makes the system's behaviour match its description and removes the missed-run noise. ## What not to do **Do not spawn a goroutine per tick** to keep the loop receiving. That replaces a self-limiting daemon with an unbounded one: runs overlap, concurrent sweeps race over the same certificates, and if the dependency slows further the in-flight count grows without limit. If you want the loop to keep ticking while a run is live, add an explicit guard — a one-slot semaphore channel or an `atomic.Bool` swapped on entry — and count the attempts you refuse, which gives you the same skipping behaviour but measured. **Do not shorten the interval.** A run that overruns will overrun a shorter interval too, and now the skipped-tick count climbs faster while nothing improves. **Do not silence the missed-run alert** because "skipping is fine for idempotent work". It usually is fine — right up until the day the run takes longer than the safety margin between the renewal window and `NotAfter`, and skipping stops being free.
- Why derive the per-run timeout from the loop's context instead of from context.Background?So cancellation still flows down. A context derived from the loop's `ctx` is cancelled both by its own deadline and by the daemon shutting down, so a run in flight ends promptly instead of holding shutdown for the rest of its budget. Deriving from `context.Background` gives you a run that keeps talking to the certificate authority after the process has been told to stop.
- What exactly goes wrong with defer cancel() written directly in the loop body?`defer` schedules the call for when the enclosing *function* returns, not when the iteration ends. In a daemon loop that function never returns, so every iteration adds another cancel function and another live child context to the parent — a slow leak that grows with uptime. Wrap the iteration's work in a function literal, or call `cancel` explicitly before continuing.
- How would you keep the loop ticking while still allowing at most one run in flight?Guard entry explicitly. A one-slot buffered channel used as a semaphore, or an `atomic.Bool` you compare-and-swap on entry and clear on exit: if the guard is taken, increment a skipped-run counter and go back to the select. You get the same skipping the ticker was already doing, but it is deliberate, measured, and visible on a dashboard.
- The run takes four minutes and the interval is one. Is widening the interval an admission of defeat?No — it is making the configuration true. The loop was already running every four minutes; the one-minute setting was fiction that generated missed-run noise. Widen it to a value the run can meet, keep the missed-run alert armed so a future regression is still visible, and treat shrinking the run as separate work with its own justification.
saying these in an interview costs you the question
- Only alerts on run errors, never on missed runs
- Writes defer cancel() straight into the loop body
- Shortens the interval when a run overruns it
- Starts a goroutine per tick with no concurrency guard
- Passes a context the work never actually checks
- Assumes the ticker will catch up once the run gets faster