skip to content

A Go tool calls setns(2) to enter a namespace, then its work sometimes runs in the wrong namespace. Why, and what is the fix?

level: seniorimportance: should knowfreq 30%

answer

  1. the kernel state belongs to the thread
  2. the goroutine did not stay put
  3. it only fails under load
  4. pin before the state change, not after
  5. let the goroutine die instead of unlocking

basics

~20 s

The setns call changes only the calling OS thread, and the Go scheduler may resume the goroutine on a different thread after a syscall or a preemption. Pin with runtime.LockOSThread before entering, and let that goroutine exit rather than unlocking.

solid answer

~50 s

Namespace membership on Linux is a property of the thread, not of the process, so `setns(2)` moves only the thread that made the call. A goroutine is not tied to a thread: after a blocking syscall, a channel operation, or a preemption it can be resumed on another thread, which is still in the original namespace, so the following work runs in the wrong place. The fix has three parts. Run the whole operation on a goroutine created for it. Call `runtime.LockOSThread()` before the `setns` call so the goroutine cannot migrate. Do not unlock and do not `defer runtime.UnlockOSThread()`; let the goroutine return so the runtime terminates that thread, because a thread left in a foreign namespace must never run other goroutines. Anything started with `go` inside the operation is unpinned and will not be in the namespace.

code

go · 13 lines
go
// enterNS wraps the setns(2) system call for one namespace fd.
func inNamespace(fd int, work func() error) error {
	errc := make(chan error, 1)
	go func() {
		runtime.LockOSThread() // must come before enterNS
		if err := enterNS(fd); err != nil {
			errc <- err
			return // still locked: the runtime kills this thread
		}
		errc <- work() // work must not start goroutines of its own
	}()
	return <-errc
}

go deeper

for a junior

Take away one rule: some operating-system state belongs to the thread, and a goroutine can change threads, so those calls need runtime.LockOSThread around them.

for a middle

Explain the ordering that matters: lock before the state change, keep every dependent call on that goroutine, and know that goroutines started inside are not pinned.

for a senior

Diagnose it from the symptoms — correct in isolation, wrong under concurrency, nothing from the race detector — and argue for exiting while locked over deferring an unlock that hands a foreign namespace to unrelated goroutines.

for a principal

Decide the shape before the bug exists. Choose between pinning a goroutine per operation, a single long-lived pinned worker, or a separate process, and bound how many pinned operations may be in flight given the process thread budget.

## The failure A tool that manages containers enters a target's namespaces per operation: open the namespace file descriptors, call `setns(2)` for each, then read or write something inside. Most of the time it works. Occasionally the read comes back from the host's namespace instead, or a mount is not visible, or a socket resolves against the wrong network stack. It gets worse under load and after apparently unrelated refactors. ## Why it happens Two facts collide. **Namespace membership is per thread.** On Linux, `setns(2)` changes the calling thread only. Same for a mount namespace created with `unshare(2)`, and for several credential and capability properties. The process as a whole has no single namespace; each of its threads has its own. **A goroutine is not a thread.** The Go runtime multiplexes goroutines over OS threads. A goroutine can be parked and resumed elsewhere whenever it blocks in a system call, performs a channel operation, is preempted, or reaches a garbage-collection safe point. So the sequence `setns` then `work` may execute the two halves on two different threads: the first thread joins the namespace, the second one never did. That is why the bug is intermittent. Nothing forces migration; a goroutine often stays on its thread for the whole operation, so tests pass. Adding a log line, a mutex, or a timeout changes the scheduling and the bug appears. ## The fix Pin the goroutine to its thread before the state change: 1. **Create a goroutine for the operation.** The pinning applies to one goroutine, and the cleanup is that goroutine's death, so it must not be a goroutine that goes on to do other work — never the caller's. 2. **Call `runtime.LockOSThread()` first**, before the `setns` call. From that point the goroutine runs only on that thread and no other goroutine runs there. 3. **Do every call that must see the namespace on that goroutine.** A helper started with `go` inside it is unpinned and runs on some other thread, outside the namespace. So is any callback you hand to a library that runs it elsewhere. 4. **Do not unlock.** Return the result over a channel and let the goroutine finish while still locked. The runtime terminates the thread, which is exactly what you want for a thread whose namespace membership you cannot reliably restore. ## Why `defer runtime.UnlockOSThread()` is the wrong reflex It looks tidier and it is worse. Unlocking returns a thread that is still joined to the target namespace to general service, where unrelated goroutines will run on it and silently inherit the namespace. The failure then moves from "my operation sometimes runs in the wrong namespace" to "unrelated parts of the program sometimes run in a container's namespace", which is far harder to attribute. Unlocking is only right when the pinned section left nothing durable behind on the thread. ## What it costs and how to bound it Every operation of this shape creates and destroys an OS thread. That is fine at tens of concurrent operations and a problem at thousands: threads are the resource that scales here, each with a kernel stack, and the runtime aborts the whole process once live threads cross its ceiling (10000 by default, settable with `runtime/debug.SetMaxThreads`). So bound how many of these operations run at once, above this code, and treat the ceiling as a tripwire rather than a limit to raise. ## Diagnosing it when you inherit it The symptom that identifies this class of bug is that the operation is correct in isolation and wrong under concurrency, with no data race reported by the race detector, because there is no shared memory involved at all. Reading the code for the pattern is faster than instrumenting it: find the state change that is documented as per-thread, then check whether a `runtime.LockOSThread` call dominates it on the same goroutine and whether anything between them can start a new goroutine. If per-thread state is being set anywhere outside a pinned, single-purpose goroutine, that is the bug. ## The alternative worth naming Some tools sidestep the whole problem by doing the namespace work in a separate process: re-execute the binary with a flag, have the child join the namespaces before the Go runtime starts many threads, and communicate over a pipe. That trades a thread for a process and removes the affinity requirement from the parent's design entirely, which is often the better answer when the pinned region would otherwise be large.

  • Why is this bug intermittent rather than deterministic?
    Nothing forces a goroutine to move. It typically keeps its thread until it blocks in a syscall, performs a channel operation, or is preempted, so a lightly loaded test run never migrates and passes. Load, an added log call, or a lock in the middle of the operation changes the scheduling and the migration starts happening.
  • Can the pinned goroutine hand the actual work to a goroutine it starts?
    No. Pinning is per goroutine and is never inherited, so the child runs on an arbitrary thread that never joined the namespace. Everything that must observe the namespace has to execute on the locked goroutine itself, which also means any library callback you use must run inline rather than on a goroutine of its own.
  • Would the race detector find this?
    No. There is no shared memory being accessed without synchronisation; the operation is entirely correct as Go code. The mismatch is between the goroutine's identity and the thread's kernel state, which the race detector does not model. Code review against the per-thread documentation is what catches it.
  • What is the alternative if the region that needs the namespace is very large?
    Move it out of the process: re-execute the binary with a flag, have the child join the namespaces at startup, and talk to it over a pipe. That trades one OS thread for one process, removes the pinning requirement from the parent's design, and makes it impossible for unrelated code to accidentally run in the namespace.

saying these in an interview costs you the question

  • Blames a data race and adds a mutex
  • Thinks setns changes the whole process
  • Adds runtime.Gosched or sleeps hoping to stay put
  • Defers UnlockOSThread and returns a polluted thread
  • Runs the real work in a child goroutine inside the lock
  • Calls LockOSThread after the namespace change