skip to content

The Unix model composes small single-purpose programs over untyped byte streams. When you design the tooling around a platform your team owns, how would you decide between that model and a single integrated program producing structured output, and what does each choice cost?

level: principalimportance: should knowfreq 20%

answer

  1. untyped streams become a contract
  2. ask who parses your output
  3. composition versus cohesion
  4. the pretty format is not an API
  5. partial failure needs structure

basics

~20 s

Decide by asking who consumes the output and how stable that contract must be. Composition over text streams wins for ad-hoc human work and independent evolution; an integrated tool with a versioned structured format wins when machines consume it and partial failure must be expressed precisely.

solid answer

~50 s

I start from the consumer. If humans at a prompt are the primary audience and the data is genuinely record-shaped, the Unix model is hard to beat: each piece evolves independently, anyone can insert a stage, and no plugin API is needed. If other programs are the audience, untyped text becomes an accidental API — the moment someone parses column four, your human-readable format is frozen and you did not agree to that. The integrated tool buys a typed, versioned contract and can express partial failure properly instead of squeezing it through one exit status; it costs you extensibility, because every new capability now goes through one codebase and one team. In practice I take the hybrid the ecosystem converged on: keep the tools small and composable, but give each one an explicit machine-readable output mode that is documented, versioned, and treated as the real contract, leaving the default format free to change.

go deeper

for a junior

Know the two shapes — several small programs joined by pipes, or one program doing the whole job — and that text passed between programs has to be parsed by whoever reads it.

for a middle

Explain concretely why parsing another tool's output is fragile, and why a documented machine-readable output mode is safer than screen-scraping the default format.

for a senior

Argue the tradeoff with real failure modes: accidental APIs, partial failure that a single exit status cannot express, and the operational cost of keeping several artefacts in version step.

for a principal

Set the policy for the platform: which outputs are supported contracts, how they are versioned and deprecated, where capabilities are allowed to live, and what the team accepts in exchange for extensibility or cohesion.

## Frame the decision as a contract question Both models are answers to the same question — *where does the coupling live?* In the composed model the coupling lives in the stream format and is implicit. In the integrated model it lives inside one program and is explicit. Neither eliminates coupling; they relocate it, and the whole judgment is about which location you can afford to defend. ## What composition actually gives you **Independent evolution.** Each program has its own release cadence, its own owner and its own tests. Adding a capability means writing a new program, not modifying an existing one — the extension point is the stream, so nobody has to design a plugin interface, and nobody's change can destabilise an unrelated feature. **Discoverable, run-time assembly.** A human at a prompt can invent a combination the authors never imagined. This is genuinely valuable during incidents, when the useful query is by definition one nobody anticipated. **Low commitment.** A small tool that turns out to be wrong is cheap to throw away. A capability inside a large tool acquires users and becomes permanent. ## What composition costs you **Accidental APIs.** The most expensive failure mode: you ship a human-readable format, someone automates against it, and now the default output is a contract you never versioned. Every cosmetic improvement becomes a breaking change discovered in someone else's pipeline. **No types, no nesting.** Byte streams model flat records well. Nested or optional data has to be flattened, and flattening is lossy or ambiguous. Values containing the delimiter are a permanent hazard. **Impoverished failure semantics.** One integer of exit status cannot say "three of four hundred records failed, here is which and why". Partial failure — the normal case in real systems — is exactly what the model expresses worst, and the default pipeline status hides upstream failure entirely unless you opt into a shell option that changes it. **Per-stage overhead and no shared state.** Each stage is a process. For high-volume or latency-sensitive paths, the process boundaries and the repeated serialise-parse cycle are real cost. **Operational surface.** Five programs means five things to install, version, sign and keep compatible with each other. That is often the deciding factor for a tool that must land on many machines. ## What the integrated tool gives you and costs you It gives you a typed internal model, precise structured errors, one artefact to install and one place to instrument. It costs you extensibility and organisational scale: every new capability queues behind one codebase's review, its test matrix grows superlinearly, and a bug in one area can take out the rest. It also tends to accumulate a poorly designed configuration language as it absorbs the flexibility that composition used to provide for free. ## The decision procedure I would actually use 1. **Name the consumers.** Human-first, machine-first, or both? Machine consumers are the strongest argument for a declared structured format, whichever packaging you choose. 2. **Ask what the failure model needs to express.** If partial success is normal and must be actionable, a single exit status is disqualifying. 3. **Ask who will extend this.** Many teams adding capabilities favours composition; one team owning a coherent workflow favours integration. 4. **Look at the data shape.** Flat records stream well; deeply nested state does not. 5. **Check the deployment cost.** How many machines, and how hard is it to keep several artefacts in step? ## The synthesis worth stating out loud The modern resolution is not to pick a side. Keep programs small and single-purpose — that part of the Unix argument has aged extremely well — but stop pretending the human format is an interface. Give every tool an explicit machine-readable output mode, document it, version it, and change the pretty format freely. Emit diagnostics on the separate error channel so structured data on the output channel stays clean. Where a stage can partially fail, put the per-item outcome *in* the structured output rather than in the exit status, and reserve the status for "did the invocation run at all". That is the honest position at this level: the Unix insight was that a universal interface enables composition, and text was simply the only universal interface available in 1975. Keep the insight and upgrade the interface.

  • How do you stop a human-readable output format from silently becoming an API?
    Publish an explicit machine-readable mode and state in the documentation that only it is supported for automation. Then actually change the human format periodically, and survey or instrument consumers early. If you never change it, you have made the promise implicitly regardless of what the docs say — the guarantee is established by behaviour, not by text.
  • Where does a single exit status genuinely fall down?
    On any operation with per-item outcomes. Processing four hundred records where three fail is the common case, and the status can only say "something went wrong". You cannot retry selectively, you cannot report precisely, and callers end up parsing the error text — which recreates the untyped-contract problem in the worst possible place.
  • Does choosing structured output mean abandoning composability?
    No — it changes the universal interface, not the principle. Small tools that read and write a common structured format still compose, and now the composition is type-checked rather than convention-checked. What you lose is the ability of any casual text utility to sit in the middle of the pipeline, which is a real cost worth naming.
  • When would you deliberately accept the accidental-API risk?
    When the tool is genuinely internal, the consumer set is small and known, and shipping speed outweighs a future migration you can coordinate in an afternoon. The mistake is making that trade for something widely distributed, where you cannot enumerate consumers and therefore can never change the format safely.

saying these in an interview costs you the question

  • Treats the Unix philosophy as a rule rather than a tradeoff
  • Assumes structured output means abandoning small tools
  • Ignores that the default output becomes an unversioned contract
  • Squeezes per-item failures through a single exit status
  • Argues from process-creation cost alone

context