skip to content

What does a "too many open files" error from net.Dial or os.Open tell you?

level: juniorimportance: must knowfreq 55%

answer

  1. the kernel is refusing to hand out more
  2. sockets, files and response bodies all count
  3. something opened is never closed
  4. defer belongs to the function, not the loop body
  5. RLIMIT_NOFILE is the ceiling

basics

~20 s

The process has hit its RLIMIT_NOFILE ceiling on open file descriptors. Every socket, file and open HTTP response body costs one, so the cause is almost always a resource opened on some code path and never closed.

solid answer

~40 s

It is the operating system refusing to give the process another file descriptor, because the process already holds as many as its `RLIMIT_NOFILE` limit allows. In Go, a descriptor is consumed by every `*os.File` from `os.Open`, every `net.Conn` from `net.Dial`, every listener, pipe and every HTTP response body still open. So the error is a symptom of accumulation, not of a bad call site: something is opened and never closed. The two classic Go causes are an error path that returns before `Close`, and `defer f.Close()` written inside a loop — `defer` runs when the *function* returns, not at the end of the iteration, so a thousand-file loop holds a thousand descriptors at once. Fix the leak first; raising the limit only moves the failure later.

code

go · 11 lines
go
func scanAll(paths []string) error {
	for _, p := range paths {
		f, err := os.Open(p)
		if err != nil {
			return err
		}
		defer f.Close() // runs at function return, not at end of iteration
		process(f)
	}
	return nil
}

go deeper

for a junior

Be ready to say what a descriptor is, name three things in Go that hold one, and show the defer-in-a-loop bug plus its fix. Interviewers ask this to see whether you close on error paths too.

for a middle

Explain that defer is registered against the function, not the block, and why an unread HTTP response body keeps a socket alive. Know that the limit is per process and has a soft and a hard value.

for a senior

Show the judgment: distinguish a count that climbs forever from one that tracks concurrency and plateaus, and say why raising the limit before diagnosing only moves the outage later in the day.

for a principal

Own the position that descriptor capacity is a stated budget, not an accident. Be able to say what number the service is designed to hold at peak and who is accountable when the code and the configured limit disagree.

## What a file descriptor is On Unix, a **file descriptor** is a small integer the kernel gives a process to refer to something it has opened. It is not only files: a regular file, a TCP socket, a pipe, a Unix socket, the runtime's own epoll/kqueue handle and the process's stdin/stdout/stderr are all descriptors. The kernel caps how many one process may hold at once with the `RLIMIT_NOFILE` resource limit, which has a **soft** limit (the one actually enforced) and a **hard** limit (the ceiling the soft limit may be raised to). When a program asks for one more than it is allowed, the kernel returns `EMFILE`, which Go surfaces in the error text as `too many open files`. (`ENFILE`, "too many open files in system", is the rarer system-wide variant.) ## What it looks like in Go The error arrives wrapped in whatever Go type the operation produces: - `os.Open` returns `*os.PathError`: `open /etc/hosts: too many open files` - `net.Dial` returns `*net.OpError`: `dial tcp 10.0.0.7:443: socket: too many open files` - an inbound `Accept` on a listener can fail the same way, so a server can also stop accepting traffic On Unix you can test for it precisely, because the error chain unwraps down to the raw errno: `errors.Is(err, syscall.EMFILE)`. Most services do not need that; the point of the error is that the *process* is out of descriptors, and the next thing it tries will probably fail too. One Go-specific detail worth knowing: on Unix the Go runtime raises the process's **soft** `RLIMIT_NOFILE` to the hard limit at startup, and `os/exec` restores the original soft limit for child processes (many C programs misbehave with a huge limit). So a Go service is usually already running at the hard limit, and "just bump the ulimit" often turns out to change nothing — the hard limit, or the container's configured limit, is the real ceiling. ## What consumes descriptors in a Go program Every one of these holds a descriptor until it is closed: - an `*os.File` from `os.Open`, `os.Create`, `os.OpenFile` - a `net.Conn` from `net.Dial`, and a `net.Listener` from `net.Listen` - an HTTP **response body** you have not closed — the connection underneath it stays open - an idle connection that `net/http`'s client keeps around for reuse The first three are yours to close. The last one is a deliberate cache, and it is bounded by settings on the HTTP transport rather than by your code. ## The two leaks that cause almost every case **1. A path that returns before Close.** Every early return, every error branch, must still close what the function opened. That is exactly what `defer` is for: put `defer f.Close()` on the line after the successful open, and no later branch can forget it. **2. `defer` inside a loop.** This is the single most common junior mistake here, and it is specific to how Go scopes `defer`. A deferred call is registered against the enclosing **function**, not the enclosing block, and runs when that function returns. So: ```go for _, p := range paths { f, err := os.Open(p) if err != nil { return err } defer f.Close() // queued — runs only when this function returns process(f) } ``` holds every file open simultaneously. Over ten files nobody notices; over fifty thousand it dies with `too many open files`. The fix is to give each iteration its own function, so each `defer` fires at the end of that call — a small named helper is clearer than an inline closure, and it also gives you a natural place to return a per-file error. ## Why you cannot lean on the garbage collector An `*os.File` does have a runtime cleanup attached that closes the descriptor once the value becomes unreachable, so a leaked file *may* eventually be closed. This is a safety net, never a strategy: collection timing is unspecified, descriptors are a kernel resource the GC does not measure or feel pressure from, and a loop can exhaust the limit long before a collection happens. Close explicitly. ## The right first move When this error appears, resist raising the limit. Ask instead whether the descriptor count is *bounded by design*: does the code hold a number that stops growing with load, or one that climbs forever? A count that climbs and never falls is a leak, and raising the limit only changes the hour at which the service dies.

  • Besides files you opened yourself, what else in a typical Go service holds descriptors?
    Every `net.Conn`, every listener, every pipe, and every HTTP response body that has not been closed — the body keeps the underlying connection alive. A client that keeps idle connections for reuse also holds descriptors on purpose; those are bounded by transport settings rather than by your code, but they still count against the same limit.
  • If an *os.File gets closed when it is garbage collected, why close it explicitly?
    Because the timing is unspecified and the collector does not track descriptors as a scarce resource. A loop can exhaust `RLIMIT_NOFILE` in milliseconds while the heap is still small enough that no collection has run. The runtime cleanup is a safety net against a crash, not a resource-management strategy — and it also means the file's buffered writes may never be checked for errors.
  • How would you confirm in Go that a specific error is the descriptor limit rather than a permissions or path problem?
    On Unix, `errors.Is(err, syscall.EMFILE)` — the `*os.PathError` or `*net.OpError` unwraps down to the raw `syscall.Errno`, so the comparison works through the wrapping. In practice you rarely branch on it; if one call site reports it, the process is out of descriptors and the next unrelated call will fail too.

File descriptors are like keys on a keyring with a fixed number of hooks. Borrowing a key is cheap; never handing one back means that eventually you cannot open anything at all.

saying these in an interview costs you the question

  • Says the fix is to raise the ulimit before finding the leak
  • Thinks defer inside a loop runs at the end of each iteration
  • Believes the garbage collector reliably closes files in time
  • Forgets that an unclosed HTTP response body holds a socket
  • Assumes only files count, not sockets or listeners