skip to content

A Go agent serving HTTP over a Unix socket cannot rebind /run/agent.sock after a hard kill. Why, and what is the fix?

level: seniorimportance: should knowfreq 36%

answer

  1. the address is a path on disk
  2. a clean Close cleans up, a kill does not
  3. who deletes the file, and when
  4. dial it before you delete it
  5. the abstract namespace has no file

basics

~20 s

A Unix listener's address is a file. Go unlinks it in the listener's Close, which a hard kill never runs, so the stale file survives and the next bind fails. Remove it only after checking nothing is live.

solid answer

~50 s

`net.Listen("unix", path)` binds by *creating* a socket file, and bind fails if anything already exists at that path — so the error is about the leftover file, not about a lingering connection state. `(*net.UnixListener).Close` unlinks the file, and that behaviour is on by default when the listener created it (`SetUnlinkOnClose` toggles it, and it is off for a listener adopted from an inherited descriptor via `net.FileListener`). SIGKILL, an OOM kill or a container stop that never reaches your shutdown path skips the Close entirely. The safe startup sequence is: `net.Dial("unix", path)` first — if something answers, another instance is live and you should abort; if nothing answers, `os.Remove` the path and listen. Deleting unconditionally is the dangerous version: it silently orphans a healthy instance, which keeps serving a socket nobody can reach. Also handle SIGTERM and close the listener so the normal path stays clean.

code

go · 12 lines
go
if c, err := net.Dial("unix", path); err == nil {
	c.Close()
	return fmt.Errorf("another instance is listening on %s", path)
}
// nothing answered, so the file is stale
os.Remove(path)

l, err := net.Listen("unix", path)
if err != nil {
	return err
}
defer l.Close() // unlinks the socket file on the clean path

go deeper

for a junior

Remember that a Unix socket's address is a real file on disk, so it can outlive the process that created it and block the next bind.

for a middle

Explain that the listener's Close performs the unlink, that SetUnlinkOnClose governs it, and why a SIGKILL therefore leaves the path occupied.

for a senior

Show the safe startup sequence and say why unconditional deletion is worse than the failure: it orphans a healthy instance. Wire SIGTERM to a clean shutdown so the hard-kill case is the only one left.

for a principal

Own the restart and identity story for local sockets: whether the supervisor owns the socket, whether the abstract namespace is acceptable, and what replaces client IP for authorisation once there is no client IP.

## The address is a file, and files outlive processes For `"tcp"` the address is a port the kernel reclaims when the process dies. For `"unix"` the address is a **path in the filesystem**, and binding it means creating a socket inode there. Nothing in the kernel removes that inode when the process exits — sockets are not like lock files with automatic cleanup. The next `net.Listen("unix", path)` calls `bind(2)`, finds an existing entry, and fails with `address already in use`. So the error means almost the opposite of what the same text means on TCP. There is no TIME_WAIT, no lingering connection, no timer that will clear it. The file will sit there until something deletes it. ## Who normally deletes it Go does, on a clean shutdown. `net.Listen("unix", path)` returns a `*net.UnixListener`, and its `Close` unlinks the socket file. That is why a program that shuts down properly leaves nothing behind and never hits this. The behaviour is controlled by `(*net.UnixListener).SetUnlinkOnClose(bool)`. The default is chosen by ownership: it is **on** when the listener created the file itself, and **off** when the listener was adopted from a file descriptor it did not create (`net.FileListener`, the shape used by socket-activation supervisors). That default is the right one — a process handed a socket by its supervisor has no business deleting the supervisor's file. What breaks the mechanism is any exit that does not run `Close`: - `SIGKILL`, including a container runtime's kill after the stop grace period - an out-of-memory kill - a host power loss - a crash path that calls `os.Exit` without running the shutdown code Any of those leaves the socket file behind, and the restart fails to bind. On a Kubernetes-style restart loop this presents as a service that came up fine the first time and now crash-loops with a confusing error. ## The startup sequence that is actually safe The naive fix — `os.Remove(path)` before every listen — trades one bug for a worse one. If a healthy instance is already serving that socket, removing the file does **not** stop it: its socket stays bound and open, it keeps serving whatever connections it has, but no new client can ever find it, because clients resolve by path. You now have a live process that appears healthy and receives no work, plus a second process bound to a new inode at the same path. That is a silent split, and it is much harder to diagnose than the bind failure you were trying to avoid. Probe first: ```go if c, err := net.Dial("unix", path); err == nil { c.Close() return fmt.Errorf("another instance is listening on %s", path) } os.Remove(path) l, err := net.Listen("unix", path) ``` A connection refused means the file is an orphan with no listener behind it — safe to remove. A successful dial means a live owner — refuse to start and let the supervisor decide. This is a small amount of code and it converts a crash-loop into either a clean start or an explicit, readable error. Two alternatives are worth knowing. On Linux the **abstract namespace** (an address beginning with `@`, for example `@metrics-agent`) has no filesystem entry at all; the name disappears when the last reference closes, so there is nothing to leave stale. The cost is that filesystem permissions no longer gate access and the name is Linux-only. The other is **socket activation**: the supervisor owns the socket, passes it as a descriptor, and your process adopts it with `net.FileListener` — in which case unlink-on-close is off and cleanup is not yours. ## Serving HTTP over it Once you hold the listener, an `http.Server` serves it exactly as it serves a TCP one — the listener is just a `net.Listener`. Two consequences are worth planning for. First, **shutdown**: closing the server closes the listener, which unlinks the file. Wire SIGTERM to a graceful shutdown and the stale-socket problem disappears from every normal restart, leaving only the hard-kill case for the startup probe to handle. Second, **there is no client address**. `http.Request.RemoteAddr` carries no useful peer identity over a Unix socket, because the peer has no network address. Anything you built on it — per-IP rate limiting, allow-lists, access logs keyed on the client address, `X-Forwarded-For` handling — stops meaning anything. Identity for a Unix-socket API has to come from somewhere else: which socket the request arrived on, an application credential in the request, or the operating system's own controls on the socket path. ## What to say in an interview Name the cause (the address is a file and a hard kill skips the unlink in `Close`), name the mechanism (`UnixListener.Close` unlinks, `SetUnlinkOnClose` controls it, off for inherited descriptors), and give the safe fix (dial before you delete), then mention the abstract namespace as the way to avoid the problem entirely on Linux.

  • Why is calling os.Remove unconditionally at startup a dangerous fix?
    It happily unlinks a socket a healthy instance is still serving. That instance keeps its bound socket and its existing connections, but no new client can reach it, because clients resolve the path. You end up with a live process receiving no work and a second process bound at the same path — a silent split that is far harder to diagnose than the bind error. Dial first and only remove an orphan.
  • What does UnixListener.SetUnlinkOnClose control, and what is the default?
    Whether closing the listener unlinks the socket file. It defaults to on when the listener created the file itself, and off when the listener was adopted from an inherited file descriptor through `net.FileListener`. That default follows ownership: a process handed a socket by its supervisor should not delete the supervisor's file when it exits.
  • When an HTTP request arrives over a Unix socket listener, what does http.Request.RemoteAddr give you?
    Nothing you can use as a client identity — the peer has no network address. Per-IP rate limiting, allow-lists and access logs keyed on the client address all stop meaning anything. Identify callers another way: which socket the request arrived on, an application credential carried in the request, or the operating system's controls on the socket path.

saying these in an interview costs you the question

  • Blames a lingering TCP TIME_WAIT state
  • Deletes the socket file unconditionally at every startup
  • Assumes the kernel removes the socket file when the process dies
  • Reaches for address-reuse options that do not apply to Unix sockets
  • Retries the bind in a loop waiting for it to clear