Unix is often described with the slogan "everything is a file". What does that actually mean for a program at the system-call level, and where does the abstraction leak?
answer
- one handle, many object types
- read, write, close on anything
- descriptors, not paths
- sockets need their own calls
- seeking fails on a pipe
basics
~20 sMost kernel objects — regular files, devices, pipes and sockets — are reached through a file descriptor and the same read, write and close calls. The abstraction leaks: sockets and terminals need their own calls, and not every object has a filesystem name.
solid answer
~50 sA file descriptor is a small integer indexing a per-process table of open objects, and the core calls — `read`, `write`, `close`, `dup`, `poll` — work on the descriptor without caring what is behind it. That is the real content of the slogan: a uniform *handle* plus a uniform byte-transfer interface. Because devices appear in the filesystem as special files, they also inherit the filesystem's ownership and permission model, so access control is one mechanism rather than several. It buys redirection, pipelines, and the ability to hand a privileged descriptor to a less-privileged process. But the abstraction is not total. `lseek` fails on a pipe or socket. Sockets are created and configured through their own calls. Terminals need their own control interface, and `ioctl` is an untyped escape hatch for everything else. And `read` may return fewer bytes than asked for on a pipe or socket, which is the classic bug.
code
c · 12 lines#include <fcntl.h>
#include <unistd.h>
#include <string.h>
int main(void) {
const char *msg = "same call, different object\n";
int fd = open("/dev/null", O_WRONLY);
write(fd, msg, strlen(msg));
write(STDOUT_FILENO, msg, strlen(msg));
close(fd);
return 0;
}go deeper
Know that a file descriptor is a small integer standing for an open thing, that 0, 1 and 2 are the standard ones, and that the same read and write calls serve files, devices and pipes.
Explain the descriptor table, how device special files reuse filesystem ownership and permissions, and name at least one concrete leak such as seeking failing on a pipe or a short read on a socket.
Show that you have been bitten: describe looping on short reads, using readiness notification for pipes and sockets, and treating an open descriptor as a capability that outlives a permission change.
Argue about the abstraction itself — what uniform handles buy an interface, where an untyped escape hatch like ioctl becomes the design smell, and how you would decide when a subsystem deserves its own call surface rather than another overloaded operation.
## What the slogan really claims The accurate version is "everything is a file *descriptor*". When a process opens something, the kernel allocates an entry in that process's descriptor table and returns its index — a small non-negative integer. From then on the process names the object by that integer, and the kernel resolves it to the underlying object. The generic operations are deliberately few: `read`, `write`, `close`, `dup`/`dup2`, `fcntl`, and the readiness calls `select` and `poll`. The second half of the claim is namespace unification: devices are exposed as **special files** in the filesystem (character devices for byte-stream hardware such as terminals, block devices for randomly addressable storage), so `open("/dev/…")` reaches hardware with the same call that reaches a regular file, and later Unix systems extended the idea further by exposing process state through a `/proc` filesystem. ## What it buys **Uniform composition.** A program that reads descriptor 0 works against a file, a terminal, a device or a pipe. Redirection and pipelines exist because of this and nothing else. **One access-control model.** A device node has an owner, a group and permission bits like any other filesystem entry, so granting a user access to a serial port is the same operation as granting access to a document. Without this, every device class would need its own authorisation scheme. **Descriptors as capabilities.** A descriptor is a live handle to an already-authorised object. A privileged process can open something, drop privilege, and keep using the descriptor; on modern Unix systems a descriptor can even be passed to another process over a local socket. Permission is checked at open time, not at every read. **Uniform inheritance.** `fork` duplicates the descriptor table, and descriptors survive `exec` unless marked close-on-exec, which is what lets a shell set up a child's environment before the child's code ever runs. ```c #include <fcntl.h> #include <unistd.h> #include <string.h> int main(void) { const char *msg = "same call, different object\n"; int fd = open("/dev/null", O_WRONLY); /* a device */ write(fd, msg, strlen(msg)); write(STDOUT_FILENO, msg, strlen(msg)); /* whatever descriptor 1 is today */ close(fd); return 0; } ``` ## Where it leaks **Not everything has a path.** A pipe created by `pipe()` has descriptors but no name in the filesystem. An ordinary network socket likewise: only a Unix-domain socket has a filesystem path. So "everything is a file" is not the same as "everything is reachable by opening a path". **The operation set is not uniform.** Sockets are born from `socket()` and then need `bind`, `listen`, `accept`, `connect`, `setsockopt` — none of which apply to a regular file. Terminals need a dedicated control interface for line discipline and terminal attributes. Where no clean call exists, `ioctl` is the catch-all: a single entry point taking an opaque request number and an untyped pointer, which is precisely the shape a uniform interface was supposed to avoid. **Positioning semantics differ.** `lseek` on a pipe, FIFO or socket fails with `ESPIPE`, because those objects have no seekable offset. Code written against regular files and then pointed at a pipe discovers this immediately. **Short reads and short writes.** On a regular file, `read` normally returns the full requested count until end of file. On a pipe, socket or terminal it returns whatever is available right now — a return of 100 for a 4096-byte request is normal, not an error, and a return of 0 means end of file. Every robust Unix program loops. This is the single most common bug caused by taking the abstraction literally. **Blocking behaviour differs.** A regular file is always "ready", so readiness notification is meaningless for it, while pipes and sockets are the whole reason `select` and `poll` exist. ## The honest summary The uniformity is at the level of *naming and lifetime* (a descriptor), and partially at the level of *byte transfer*. It was never total, and Plan 9 — designed by the same group at Bell Labs — was in part an argument that Unix had not taken its own idea far enough. In an interview, the answer that lands is: name the descriptor as the unifying concept, give the concrete win (one permission model, redirection, pipelines), then volunteer a leak such as `ESPIPE` on a pipe or short reads on a socket. That combination shows you have written the code rather than read the slogan.
- If descriptors unify everything, why does ioctl exist at all?Because device-specific operations do not fit read and write. Setting a terminal's baud rate or querying a device's geometry transfers control information, not a byte stream. `ioctl` is the deliberate escape hatch: one entry point carrying an opaque request code and an untyped pointer. It preserves the descriptor as the handle while admitting that the *operations* were never fully uniform.
- What is the practical consequence of a descriptor being a capability rather than a name?Permission is checked once, at open time. A process can open a resource while privileged, drop privilege, and keep working through the descriptor — and it can pass that descriptor to another process over a local socket, granting access without granting the ability to reopen the path. It also means revoking filesystem permissions does not close descriptors that are already open.
- Why does a program that works on files sometimes truncate data when its input is a pipe?Because it assumed `read` fills the buffer. On a regular file that is nearly always true; on a pipe, socket or terminal `read` returns whatever has arrived so far. Code that treats a short return as end of input drops the remainder. The fix is to loop until the call returns 0 for end of file or a genuine error.
A file descriptor is like a coat-check ticket: the number tells you nothing about the coat, and you hand it back to get the same service regardless of what is hanging on the hook — but some items still need special handling at the counter.
saying these in an interview costs you the question
- Claims every kernel object has a path in the filesystem
- Says sockets are used only through read and write
- Treats a short read as an error or as end of file
- Assumes lseek works on any descriptor
- Thinks the slogan means files are stored as plain text