What are CPU affinity and per-core run queues, and when is deliberately pinning work to specific cores worth the loss of scheduler flexibility?
answer
- per-core queues avoid a global lock; stealing rebalances
- soft affinity = prefer last core (cache still warm)
- pinning buys cache + memory locality + predictability
- costs flexibility, portability, topology maintenance
- pin few threads + isolate the cores; derive map at runtime
basics
~20 sSchedulers keep per-core run queues and prefer to resume a thread on its previous core, because its cached data is there; idle cores steal work from busy queues. Explicit pinning fixes a thread to chosen cores, protecting cache and memory locality for latency-critical work, at the cost of the scheduler's ability to balance load.
solid answer
~60 sModern schedulers avoid a single global run queue — it would be a contention point at every scheduling decision. Instead each core has its own queue; idle cores **steal** work from busier ones. Schedulers also apply **soft affinity**: they prefer to resume a thread on the core it last ran on, because its data is still in that core's caches, and they resist migration unless the imbalance justifies it. **Hard affinity** (pinning) is you overriding that: this thread may run only on these cores. It buys cache warmth, predictable placement, and on multi-socket machines it keeps a thread near its own memory instead of paying remote-access latency. Pinning plus isolating those cores from other work is how low-latency systems cut tail jitter. The cost is real: a pinned thread cannot use an idle core elsewhere, so a hot spot cannot be balanced away, and a bad map hurts more than no map. Justify it with measured tail latency, not intuition — pin only the few threads whose jitter you are buying down.
code
text · 9 linescore0 [T1 T2 T3] core1 [T4] core2 [] core3 [T5 T6]
core2 goes idle:
1. check own queue -> empty
2. steal from near neighbour (shares cache with core3) -> take T6
3. only if that fails, steal across socket (colder, remote memory)
soft affinity: when T1 blocks and wakes, prefer core0 again
(its data is still in core0's caches)go deeper
Know that affinity means restricting which cores a thread may run on, and that schedulers already prefer to reuse a thread's previous core for cache reasons.
Explain per-core run queues and work stealing, and why migration is costly — the thread's cached working set stays behind on the old core.
Justify pinning from measured migration and tail-latency evidence, pin narrowly, isolate the cores, and derive the map from runtime topology.
Treat it as an architectural commitment — shard-per-core designs, interrupt steering, memory locality on multi-socket hardware, portability across container CPU sets — and be explicit about the operational ownership it creates and when to remove it.
## Why per-core run queues exist A single global list of runnable threads is conceptually simple and practically bad: every core touches it on every scheduling decision, so it becomes a lock and a cache-line hot spot that gets worse with more cores. The standard answer is a **per-core run queue**. Each core makes its own decisions locally with no shared lock in the common path. That introduces imbalance — one queue may be long while another is empty — which is solved by **work stealing** (or periodic balancing): an idle core takes runnable work from a busy core's queue. Stealing is deliberately biased: cores prefer to steal from topologically near neighbours (same cache-sharing cluster, then same socket) before reaching far, because pulling a thread across the machine throws away its cached state. ## Soft affinity Schedulers already implement a weak form of affinity without being asked. When a thread becomes runnable again, resuming it on the core it last used means its data may still be in that core's private caches and its translation entries may still be present. So the scheduler prefers the previous core and requires a meaningful imbalance before migrating a thread away. "Cache affinity" is this preference; it is a heuristic, not a guarantee. ## Hard affinity: pinning Pinning restricts a thread (or process) to an explicit set of cores. What it buys: 1. **Cache locality you can rely on.** A pinned thread's working set stays in one core's cache hierarchy instead of being rebuilt after every migration. 2. **Memory locality on multi-socket machines.** On a non-uniform memory architecture, memory attached to another socket costs meaningfully more per access. Pinning a thread near the memory it allocated keeps accesses local; migrating it silently converts every access into a remote one. 3. **Predictability.** Combined with isolating cores from general-purpose work, pinning removes the interference that produces tail-latency jitter — an interrupt or a background task landing on your hot core. 4. **Structural designs.** Shard-per-core architectures assign each shard its own core and its own data, so there is no cross-core sharing and often no locking at all in the data path. Pinning is what makes that model real rather than aspirational. ## What pinning costs - **Lost flexibility.** The scheduler can no longer move a pinned thread to an idle core. If your assignment is uneven, cores sit idle while pinned work queues — and the scheduler is forbidden from fixing it. - **Brittleness across environments.** Core counts, topology, hyper-threading layout and container CPU sets differ between development, staging and production. A hard-coded map that was optimal on one machine can be pathological on another, and a pinned set that does not exist in a restricted cpuset is a startup failure or a silent collapse onto one core. - **Interference with sibling threads.** Pinning to logical cores that are hardware threads of the same physical core gives you two threads competing for one core's execution resources — often worse than being spread out. - **Operational weight.** Someone must own the map, keep it correct as the topology or deployment changes, and understand it during an incident. ## Deciding A defensible decision procedure: 1. **Start with none.** The scheduler's soft affinity and topology-aware stealing are good, and they adapt. 2. **Measure the specific symptom.** Pinning targets *jitter and migration*, not average throughput. The evidence you want is high migration counts on hot threads, tail latency far above the median, or per-access cost that changes with placement. If the median is your problem, pinning is the wrong tool. 3. **Pin narrowly.** A handful of latency-critical threads — a polling network thread, a per-shard worker — not the whole application. Everything else stays under the scheduler. 4. **Isolate as well as pin.** Pinning your thread to a core that others also use gains little; the benefit comes from that core being reserved. That means excluding general work and, where relevant, steering device interrupts away. 5. **Derive the map at runtime** from the topology and the effective CPU set the process was given, never hard-coded constants. Fail loudly if the expected topology is absent. 6. **Re-measure after every hardware, container or topology change**, and be willing to remove it: an affinity map that no longer matches reality is worse than none. ## The honest summary Affinity is a specialization tool. For the overwhelming majority of services, letting the scheduler balance is right, and effort is better spent on reducing runnable-thread count and work per request. For a narrow class — packet processing, market data, storage engines, shard-per-core designs — pinning plus core isolation is the difference between a controlled tail and an uncontrolled one. The senior signal is knowing which situation you are in and holding measured evidence for the choice, not treating pinning as a general performance technique.
- Why do schedulers use per-core run queues instead of one global queue?A global queue is touched by every core on every scheduling decision, making it a lock and cache-line hot spot whose cost grows with core count. Per-core queues let each core decide locally with no shared lock on the fast path. The resulting imbalance is corrected by work stealing, which is biased toward topologically near cores so migrations lose as little cache warmth as possible.
- What makes pinning especially valuable on a multi-socket machine, and especially dangerous to get wrong?Memory is attached to sockets, so an access to memory on another socket costs noticeably more than a local one. Pinning a thread near the memory it allocated keeps its accesses local and its latency stable. Get the map wrong — thread on one socket, its data on another — and you convert every memory access into a remote one, which can be worse than no pinning at all.
- What evidence would convince you that pinning is the right intervention?Measured migration counts on the hot threads, tail latency far in excess of the median with a placement-correlated pattern, and an experiment showing the tail improves when the threads are pinned and their cores isolated. If only the median or aggregate throughput is unsatisfactory, the problem is work volume or concurrency level, and pinning will not address it.
Soft affinity is sending a worker back to the bench where their tools are already laid out. Pinning is nailing them to that bench: fast when the work is theirs, useless when the next-door bench is idle and the queue is at their feet.
saying these in an interview costs you the question
- Treating pinning as a general throughput optimization rather than a jitter and locality tool.
- Ignoring that pinning prevents the scheduler from using idle cores elsewhere.
- Hard-coding core numbers instead of deriving the map from the runtime topology and the process's allowed CPU set.
- Pinning to two logical cores that are hardware threads of the same physical core and expecting isolation.
- Assuming a global run queue is how schedulers work, and forgetting work stealing and soft affinity already exist.