skip to content

In Go's runtime, what preempts a goroutine that has run too long, and after roughly how long?

level: middleimportance: should knowfreq 32%

answer

  1. nothing hands out fixed time slices here
  2. a background loop watches each processor
  3. it compares counters between two visits
  4. the threshold is a round number of milliseconds
  5. it acts at 10 ms of uninterrupted running

basics

~20 s

sysmon, the runtime's background monitor, checks each processor periodically and flags any goroutine that has held one for more than about 10 milliseconds, so the scheduler preempts it. A stop-the-world request preempts every running goroutine at once.

solid answer

~50 s

Go's goroutine scheduler has no timer interrupt handing out fixed slices. Instead sysmon, a runtime monitor loop, wakes periodically — its interval backs off from tens of microseconds up to 10 ms — and compares each processor's scheduling tick with what it saw on the previous visit. If the same goroutine has held that processor for longer than 10 ms, sysmon marks it for preemption: the goroutine's stack limit is poisoned so its next function call yields, and since Go 1.14 the thread is signalled too, so a call-free loop is interrupted as well. The other trigger is a stop-the-world request, which asks every running goroutine to stop immediately. So 10 ms is a threshold for forcible preemption, not a guaranteed time slice — the real delay also depends on when sysmon next looks and when the goroutine can be safely stopped.

go deeper

for a junior

Remember that Go's runtime watches for goroutines that run a long time and preempts them, so no single goroutine can hold a processor forever. The rough figure is 10 milliseconds.

for a middle

Explain how the decision is made: a monitor loop compares each processor's scheduling counter between visits and flags a goroutine that has not changed for more than 10 ms, then the ordinary preemption mechanisms do the stopping.

for a senior

Be able to argue why 10 ms is a fairness bound rather than a slice, and why the urgent path - everyone stops at once - is where preemption latency actually shows up in your pause times.

for a principal

Own the consequence for workload placement: a threshold-based fairness mechanism means CPU-heavy goroutines and latency-critical ones in one process interact, and the policy call is whether they should share a process at all.

## The absence of a timer interrupt An operating-system scheduler is driven by a hardware timer: the clock fires, the kernel takes the CPU back, and it reassigns it. Go's goroutine scheduler has no such thing. Goroutines are user-space objects multiplexed onto OS threads; the kernel has no idea they exist and will never interrupt one on the runtime's behalf. Everything about taking a goroutine off a processor has to be organised by the runtime itself. Most of the time no organisation is needed, because goroutines stop themselves constantly: a channel receive, a mutex wait, a network read, a sleep. Well-behaved concurrent Go code yields far more often than any timer would force it to. What is left is the awkward case: a goroutine that simply keeps computing. ## sysmon, the monitor loop The runtime keeps a monitoring loop called sysmon. It wakes up, looks around, and goes back to sleep. Its sleep interval is adaptive — it starts in the tens of microseconds and backs off toward 10 ms when the program is idle, so the loop is cheap on a quiet process and attentive on a busy one. Among its jobs is deciding whether a goroutine has monopolised a processor. It cannot ask directly, so it uses a counter: each processor records how many times it has scheduled a goroutine. sysmon remembers the counter and the timestamp from its last visit. If the counter has not moved since then and more than 10 milliseconds have passed, the same goroutine is still running, and sysmon requests its preemption. That request is not itself a context switch. It is a flag, delivered through the two mechanisms the runtime has: the goroutine's stack limit is poisoned so that the next function call traps into the runtime, and — since Go 1.14 — a signal is sent to the OS thread so a loop containing no calls can be interrupted as well. ## Why 10 ms is a threshold, not a slice It is tempting to translate '10 ms' into 'goroutines get a 10 ms time slice'. That is wrong in both directions: - **Most goroutines never reach it.** A request handler that reads from a socket and waits on a database yields within microseconds. It is descheduled by blocking, not by sysmon, and 10 ms never enters the picture. - **A goroutine can run longer than 10 ms.** The clock only starts mattering when sysmon next wakes and looks; then the goroutine must reach a point where it can be safely stopped. Both add latency after the threshold has passed. So the number is a bound on unfairness, not a quantum. It is the runtime saying: no goroutine should be able to hold a processor indefinitely. ## The other trigger: everyone stops at once sysmon handles the slow, fairness-driven case. The urgent case is an operation that requires every goroutine to pause — a stop-the-world. There, the runtime does not wait for a 10 ms threshold: it requests preemption of every running goroutine immediately and waits for each to comply. This is why preemption latency has consequences beyond fairness. If one goroutine takes a long time to reach a point where it can be stopped, every goroutine that has already stopped stays stopped, waiting for it. Before Go 1.14, when a call-free loop could not be stopped at all, this was catastrophic: the pause request never completed. With asynchronous preemption the wait is bounded and normally sub-millisecond, but the shape of the problem — the slowest goroutine to stop sets the pause length — remains the reason preemption is treated as a latency mechanism, not just a fairness one. ## How to talk about it A good answer names the actor, the signal it watches, and the number: sysmon compares each processor's scheduling tick between visits, and preempts a goroutine that has held one for more than 10 ms; it does so by flagging the goroutine so its next call yields and by signalling the thread if it makes no calls. Then it adds the qualifier that separates a memorised fact from understanding: 10 ms is a ceiling on how long a goroutine can hog a processor, not a slice each goroutine is handed, and a stop-the-world request preempts everyone without waiting for it.

  • Does this mean each goroutine gets a 10 ms time slice?
    No. It is a threshold for forcible preemption, not a quantum. Most goroutines yield within microseconds by blocking on a channel, a lock or I/O and never come near it. And a goroutine can exceed 10 ms, because the clock only bites when sysmon next looks and when the goroutine can be stopped safely.
  • What besides sysmon causes every running goroutine to be preempted at once?
    A stop-the-world request. When the runtime needs every goroutine paused, it asks all of them to stop immediately rather than waiting on a 10 ms threshold, and waits for the last one to comply. That is why the slowest goroutine to reach a safe point sets how long the pause lasts.
  • Why can't Go rely on the operating system's timer interrupt to schedule goroutines?
    Because the OS schedules threads and knows nothing about goroutines; many goroutines share each thread. Beyond that, the runtime must stop a goroutine at an instruction whose stack and registers it can describe for the collector, which the kernel cannot arrange. Hence safe points, sysmon, and the runtime's own preemption signal.

saying these in an interview costs you the question

  • Says every goroutine receives a fixed 10 ms time slice
  • Claims the OS scheduler preempts individual goroutines
  • Thinks a goroutine can only be descheduled when it blocks
  • Says sysmon stops the goroutine directly rather than flagging it
  • Treats a function call as an automatic goroutine switch