The Unix philosophy is usually summarised as "write programs that do one thing well, and write programs to work together". What concrete design conventions make that composition possible, and what does the approach cost you?
answer
- one job per program
- streams in, streams out
- the shell wires, the kernel buffers
- text carries no schema
- last command sets the status
basics
~20 sUnix builds small single-purpose programs that read bytes from standard input and write bytes to standard output, so a shell can chain them with pipes. Composition becomes nearly free; structured data and precise error reporting do not.
solid answer
~50 sThree conventions carry the whole idea. First, every process starts with three inherited file descriptors — standard input, standard output and standard error — so a program reads and writes without knowing whether the other end is a terminal, a file or another program. Second, the kernel provides pipes, and the shell wires one program's output descriptor to the next one's input descriptor, so composition is decided at run time rather than designed into either program. Third, a program reports success or failure as a small integer exit status, which lets the shell branch on it. Diagnostics go to standard error so they never contaminate the data stream. The payoff is that tools written decades apart still interoperate. The cost is that a line of text carries no schema: consumers parse by convention and break on spaces, changed columns or locale differences, and an exit status is a very thin error channel.
code
bash · 1 linegrep ERROR app.log | cut -d' ' -f1 | sort | uniq -c | sort -rn | head -5go deeper
Be able to say what standard input, standard output and standard error are, and show a two- or three-stage pipeline you have actually used. Knowing that exit status zero means success is expected.
Explain the mechanics: the shell forks and rewires descriptors before exec, the pipe is a kernel buffer that blocks in both directions, and the stages run concurrently. Name at least one concrete parsing hazard.
Show judgment about when to stop composing. Argue for a stable machine-readable output contract, describe how you handle failures that a single exit status cannot express, and point at where per-process overhead makes a pipeline the wrong shape.
Own the tradeoff at the platform level: which interfaces in your systems are stable contracts versus convenience formats, how you version them, and what you accept when you choose human-readable text as an integration surface.
## The rule Doug McIlroy's much-quoted summary from Bell Labs is: make each program do one thing well; expect the output of every program to become the input to another, as yet unknown, program; and write programs to handle text streams, because that is a universal interface. It is a statement about **interfaces**, not about program size — a tool is "one thing" when its contract is one thing. ## The three conventions that make it work **Standard descriptors.** A Unix process starts life with file descriptor 0 open for input, 1 for output and 2 for diagnostics, inherited from whatever started it. A program that reads 0 and writes 1 is automatically usable against a keyboard, a file, a device or another process; it never contains code to decide which. That is why redirection needs no cooperation from the program. **Pipes.** The `pipe()` system call returns a pair of descriptors joined by an in-kernel buffer. To build `a | b`, the shell creates the pipe, forks twice, points the left child's descriptor 1 at the write end and the right child's descriptor 0 at the read end, then execs both. Both processes run *concurrently*; the kernel buffer supplies backpressure — the writer blocks when the buffer is full, the reader blocks when it is empty, and the reader observes end-of-file only when every write end has been closed. Nothing is staged through a temporary file. ```sh grep ERROR app.log | cut -d' ' -f1 | sort | uniq -c | sort -rn ``` Five independent programs, none of which knows the others exist, produce a ranked error report. The composition lives in the command line, not in any of the programs. **Exit status.** A process returns an integer to its parent, zero meaning success. That single convention is what makes `&&`, `||` and scripted control flow possible across programs written in different languages. **Separate diagnostic channel.** Because errors go to descriptor 2, a warning printed mid-run does not appear as a data record to the next stage of a pipeline. ## What the model buys Composition is combinatorial rather than planned: *n* small tools yield far more useful combinations than *n* features inside one program, and no tool needs a plugin API to be extended, because the extension point is the stream itself. Each tool is independently testable, independently replaceable, and can be written in any language that can read and write bytes. And because the interface is bytes, the wiring can be decided interactively by a human at a prompt — the reason ad-hoc analysis on a Unix box is so fast. ## What the model costs **No schema.** "Whitespace-separated columns of text" is a convention, not a contract. A filename containing a space, a newline in a field, a locale that formats dates or sorts characters differently, or a tool that adds a column in a new release will all break a downstream parser silently. The NUL-delimited convention (`find -print0` feeding `xargs -0`) exists precisely because the default delimiter is unsafe — and note that neither of those options is in POSIX; they are widely implemented extensions. **Thin error semantics.** One integer plus free-form prose on standard error is a poor substitute for a structured failure value. There is no portable way to say *which* record failed and why. **Failure hiding in pipelines.** A POSIX shell reports a pipeline's exit status as the status of the *last* command, so a failure upstream is discarded unless you use a shell option such as `pipefail` in bash, ksh or zsh. **Boundary cost.** Every stage is a separate process with its own creation cost, its own memory, and no shared type system. For tight loops over huge data, one program doing three transformations beats three programs doing one each. **Poor fit for nested data.** Line-oriented streams model records well and trees badly, which is why modern tools bolt structured output formats back on top of the same descriptor interface. ## Why interviewers ask it They are checking whether you can reason about interface design, not whether you can quote McIlroy. The strong answer names the mechanisms (inherited descriptors, kernel pipes, exit status), states the payoff honestly, and then volunteers the failure modes — because the same tradeoff reappears every time you decide between a stable machine-readable contract and a convenient human-readable one.
- In that pipeline, do the programs run one after another or at the same time?At the same time. The shell forks every stage and connects them with kernel pipes before any of them runs, so data flows through as it is produced. The kernel buffer provides flow control: a fast producer blocks once the buffer fills, and a fast consumer blocks waiting for data. Nothing accumulates the full output of one stage before the next starts.
- Why is a filename containing a space such a recurring source of breakage in this model?Because whitespace is the default field separator, so a name with a space parses as two fields. The tools are not wrong — the stream simply has no way to say where one value ends. The usual mitigation is a delimiter that cannot occur in a filename, which is why NUL-separated output exists as an option on find and xargs, though it is an extension rather than part of POSIX.
- Does "do one thing well" mean a program must stay small?No — it means its contract should be one thing. A compiler is enormous and still obeys the philosophy, because it exposes one job through one predictable interface. The rule is violated by a tool that mixes unrelated responsibilities behind one entry point, not by a tool that is internally large.
saying these in an interview costs you the question
- Says the shell buffers all output before starting the next stage
- Claims a pipeline fails as soon as any stage fails
- Treats "do one thing" as a rule about lines of code
- Assumes columns in a tool's output are a stable contract
- Thinks pipes cannot carry binary data