What actually happens when the operating system switches from running one thread to another, and where does the cost of that switch come from?
answer
- save registers, pick next, restore, maybe swap page tables
- direct cost small; cache re-warm dominates
- cross-process adds address-space switch + translation entries
- voluntary (blocked) vs involuntary (preempted)
- migration = caches on the wrong core
basics
~20 sThe kernel saves the running thread's registers and program counter, picks another runnable thread, and restores its state; a switch to a different process also swaps page tables. The direct cost is small. The larger, indirect cost is cold caches, branch predictors and address-translation entries when the new thread starts running.
solid answer
~50 sMechanically: an interrupt or system call enters the kernel, the outgoing thread's registers and program counter are saved into its control block, the scheduler picks the next runnable thread, its state is restored, and execution resumes at its saved point. If the next thread belongs to a different process, the address space is swapped too — new page tables, and address-translation cache entries that may be invalidated. That direct work is measured in microseconds or less. The **indirect** cost usually dominates: the incoming thread starts with caches full of someone else's data, so it stalls on memory until its working set is re-fetched, and branch predictors and prefetchers must relearn. The more the two threads' working sets differ, the worse it is. Hence the practical rules: switches within a process are cheaper than across processes; switching more often than your working sets can be re-warmed wastes cycles; and pinning or grouping related work reduces the damage.
code
text · 8 linestimer interrupt
-> enter kernel mode
-> save T1: registers, PC, SP, flags (vector state lazily)
-> scheduler picks T2 from run queue
-> if T2 in another process: install T2's page tables
-> restore T2 context, return to user mode
-> T2 resumes ... and now stalls on cache misses for its own data
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ this part is usually the expensive halfgo deeper
Describe the mechanics: save the current thread's registers and program counter, choose another thread, restore its state and continue.
Separate direct cost from indirect cost, explain why cross-process switches add an address-space change, and why cold caches usually dominate.
Use the voluntary/involuntary split to diagnose real systems, and connect switch cost to spin-versus-park decisions and to keeping runnable threads near core count.
Reason about the memory hierarchy and topology — migration, cache-sharing groups, remote memory on multi-socket machines — and how those shape placement and concurrency limits.
## The mechanics A context switch is triggered by one of: a timer interrupt ending a time slice, a thread blocking (I/O, lock, queue wait), a thread yielding, a higher-priority thread becoming runnable, or a device interrupt. The kernel then: 1. **Enters kernel mode** — a mode transition with its own cost. 2. **Saves the outgoing thread's context** — general registers, program counter, stack pointer, status flags, and possibly floating-point/vector state (often saved lazily, because that block is large and many threads never touch it). 3. **Updates bookkeeping** — accounting, run-queue position, state transition from running to runnable or blocked. 4. **Chooses the next thread** from a run queue. 5. **Restores that thread's context** and resumes it at its saved program counter. If the incoming thread lives in a **different process**, add an address-space switch: install different page tables. Historically that invalidated the entire translation cache (the TLB, which caches virtual-to-physical mappings); modern CPUs tag entries with an address-space identifier so both processes' entries can coexist, which softens but does not remove the effect. ## Direct versus indirect cost **Direct cost** is the instructions above: mode transitions, register save/restore, scheduler bookkeeping. It is small and fairly constant — on the order of a microsecond, often less. **Indirect cost** is what happens *after* the switch. The new thread runs with: - **Cold data caches.** Its working set was evicted while it was descheduled. Every miss that goes to main memory costs hundreds of cycles. - **Cold instruction cache** for a different code path. - **Cold branch predictors and prefetchers**, so early execution mispredicts more. - **Possibly cold translation entries**, adding page-table walks on top of the cache misses. This is why measured switch cost varies so widely: it depends on how much of the incoming thread's working set survived, which depends on how long it was away and what ran in between. Two threads working on the same data on the same core may switch almost for free; two threads with megabyte working sets on the same core can cost far more in re-warming than in the switch itself. ## Voluntary versus involuntary A **voluntary** switch happens because a thread blocked — it had nothing to do anyway, so the switch is mostly recovering otherwise-idle time. An **involuntary** switch happens because the scheduler preempted a thread that still had work. High involuntary-switch counts are the interesting signal: they mean more runnable threads than cores, and time going into re-warming rather than progress. The distinction matters when diagnosing: a service with many voluntary switches is probably I/O-bound and fine; the same service with many involuntary switches is contending for CPU. ## Related but distinct costs - **Mode switch (system call)** — user to kernel and back, *without* changing threads. Cheaper than a context switch; often confused with it. - **Blocking on a contended lock** — the direct cost is a switch, but the real cost is the lost parallelism and the convoy of waiters behind it. This is why brief contention is often better handled by spinning a short while before parking: if the wait is shorter than a switch plus re-warm, spinning wins. - **Thread migration** — being resumed on a *different* core than before, so even caches that survived are on the wrong core. On multi-socket machines, migrating across sockets can also mean the thread's memory is now remote, adding latency to every access. ## What follows practically 1. **Fewer, longer runs beat many short ones** for throughput; a scheduler quantum exists precisely to amortize the switch over useful work. 2. **Keep runnable threads near the number of cores** for CPU-bound work; each extra runnable thread buys no throughput but adds switches and cache pressure. 3. **Group related work.** Threads sharing data benefit from running on the same core or the same cache-sharing group, which is the rationale behind affinity and behind per-core work queues. 4. **Prefer parking over spinning for long waits, spinning for very short ones** — the crossover is roughly the switch plus re-warm cost. 5. **Measure switches per second, split voluntary and involuntary**, before concluding that switching is your problem. Very often it is a symptom of oversubscription rather than the root cause. The headline to carry into an interview: a context switch is cheap to perform and expensive to recover from, and almost all of the expense is memory-hierarchy state you cannot see in a syscall trace.
- Why can a context switch between two threads of one process be cheaper than between two processes?Same-process threads share the address space, so no page-table swap is needed and translation-cache entries stay valid. They also often share code and some data, so more of the incoming thread's working set is still in cache. A cross-process switch changes the mapping and typically leaves colder caches and more translation misses.
- When is spinning on a lock preferable to blocking, in switch-cost terms?When the expected wait is shorter than the cost of a context switch plus re-warming the caches — typically very short critical sections on a machine with a spare core. If the wait may be long, spinning burns a core that another runnable thread could have used, so parking wins. Adaptive strategies spin briefly, then park.
- Your monitoring shows a jump in context switches per second. What do you check before concluding the switches are the problem?Split voluntary from involuntary. Voluntary switches usually mean threads are blocking on I/O or locks — the switch is a symptom, and the wait is the story. A rise in involuntary switches means more runnable threads than cores, so the fix is reducing concurrency or CPU work, not micro-optimizing the switch itself.
Swapping the person at a workbench takes a moment. The real delay is that the newcomer's tools and half-finished parts were cleared away and must be fetched back.
saying these in an interview costs you the question
- Quoting a single universal cost for a context switch as if it were a constant.
- Counting only register save/restore and ignoring cache, branch-predictor and translation re-warming.
- Confusing a system call (mode switch) with a thread switch.
- Assuming more threads always means more throughput, so extra switching is harmless.
- Ignoring the difference between voluntary and involuntary switches when diagnosing.