skip to content

In bash, `unzip <(curl -sL "$url")` fails to read the archive, although saving the same URL to a file and unzipping it works. Given what `<(...)` hands a command, why does unzip fail here while `diff <(...)` is perfectly happy?

level: seniorimportance: nice to knowfreq 20%

answer

  1. a name, but not a file
  2. one pass, forward only
  3. the index lives at the end of a zip
  4. lseek on a pipe returns ESPIPE

basics

~20 s

The /dev/fd path from <(...) names a pipe, which is readable once and forward only. unzip must seek to the central directory at the end of the archive before it can read entries, and a pipe cannot be seeked; diff and grep only stream forward, so the same path suits them fine.

solid answer

~50 s

Process substitution gives you a path, but not a *file*. `/dev/fd/63` names the read end of a pipe, so the consumer gets one forward-only pass: no `lseek`, no rewind, no second open that starts from the beginning. `diff`, `grep`, `sort` and `wc` are all happy with that because they read straight through. The ZIP format puts its central directory at the *end* of the archive, so `unzip` seeks there first and then seeks back to each entry — it fails immediately on a pipe, which is exactly why `unzip` cannot read from stdin either. The same limitation catches anything that mmaps or index-seeks its input, such as SQLite database files or some media containers. When the consumer needs a real file, materialise one: `tmp=$(mktemp) && curl -sL "$url" -o "$tmp"` with cleanup on exit. Note that bash's FIFO fallback on systems lacking `/dev/fd` is not seekable either.

go deeper

for a junior

Take away the shape of the rule: <(cmd) gives a command a stream with a filename, so tools that read straight through are fine and tools that jump around inside a file are not.

for a middle

Name the mechanism — the path is the read end of a pipe, so lseek fails with ESPIPE and there is no size, no rewind, and no second pass — and match that against how a consumer reads.

for a senior

Diagnose it from the error rather than by trial and error: a tool complaining about structure or a missing index on a /dev/fd path means it wanted seekability. Then reach for mktemp with cleanup on exit, or change the artifact format to a streamable one.

for a principal

Generalise the constraint when reviewing pipelines: process substitution removes temp-file lifecycle management, never the requirement for file semantics. Publishing streamable artifacts is usually the cheaper systemic fix than working around ZIP in every consumer.

## The path is a pipe wearing a filename `<(cmd)` creates a pipe, runs `cmd` with its stdout on the write end, and substitutes a path naming the read end — `/dev/fd/63` on Linux and macOS. The path resolves, `open()` succeeds, `read()` returns data. What does *not* work is everything that assumes a file has a size and a position you can move around in: - `lseek()` on a pipe fails with `ESPIPE`. - `fstat()` reports a FIFO, with no meaningful size. - `mmap()` is not possible. - The data is consumed as it is read; there is no rewinding to byte 0. So the correct mental model is "a stream that happens to have a name", and the question for any consumer becomes: does this program read forward, or does it navigate? ## Why ZIP in particular The ZIP container stores its index — the central directory — at the *end* of the file, with a locator record in the last bytes. A reader must therefore seek to the tail, parse the directory, and then seek backwards to each member's local header. That design is why `unzip` has never had a stdin mode, and it is why `unzip /dev/fd/63` fails with an error about not finding the zipfile directory rather than a permissions or missing-file error. The failure has nothing to do with `curl`, the URL, or bash's syntax. Contrast the formats built for streams. `tar` reads member headers inline as it goes, so `tar -xzf <(curl -sL "$url")` genuinely works, as does `curl … | tar -xz`. The difference is the archive format, not the shell. ## Which tools break As a rule of thumb, a consumer breaks on a substituted path when it: - reads an index stored at the end (ZIP, and anything else with a trailer); - opens the file more than once, or re-reads after EOF; - memory-maps it (SQLite opening a database file, for instance, needs a real seekable file — a database on `/dev/fd/63` is not usable); - needs the size up front to allocate or to report progress; - randomly accesses records by offset. And it works fine when the consumer is a straight-through reader: `grep`, `sed`, `awk`, `sort`, `wc`, `diff`, `comm`, `join`, `gzip`, `tar`. ## The FIFO fallback does not rescue you On a system with no `/dev/fd`, bash implements process substitution with a real named pipe created in a temporary directory. That path exists on the filesystem, which tempts people to think it behaves like a file — it does not. A FIFO has no contents and no seekable position; it is the same stream with a different name. Nothing about the fallback makes `unzip` work. ## The other ways the path fails to resolve Two more failures wear the same disguise and are worth recognising, because both produce "no such file or directory" on a path you can see printed: **Across a machine boundary.** `ssh host unzip <(curl -sL "$url")` expands locally. The remote shell receives the literal text `unzip /dev/fd/63` and opens *its own* descriptor 63 — an unrelated object or nothing at all. **Across `sudo`.** `sudo` closes inherited file descriptors above stderr before executing the target command, so by the time the program opens `/dev/fd/63` there is no such descriptor in its table. `sudo tool <(gen_config)` therefore fails even though the same line without `sudo` works. ## What to do instead When the consumer needs a real file, give it one and take responsibility for the cleanup: ```bash tmp=$(mktemp) trap 'rm -f "$tmp"' EXIT curl -fsSL "$url" -o "$tmp" unzip -q "$tmp" -d ./out ``` That is three extra lines and it is the right answer — `-f` so a 404 is an error rather than an HTML file, cleanup on every exit path, and a genuine seekable file for `unzip`. If you would rather stay streaming, change the format instead of the plumbing: publish a `.tar.gz` and pipe it. The general lesson generalises past bash: process substitution removes the *temp-file management* problem, not the *file semantics* requirement. If the program on the other end truly needs a file, no amount of shell cleverness will conjure one.

  • Why does `tar -xzf <(curl -sL "$url")` work when the unzip version does not?
    The tar format is stream-oriented: each member is preceded inline by its own header, so a reader can extract everything in one forward pass without ever knowing what comes later. ZIP puts its central directory at the end of the archive, forcing a seek to the tail before any entry can be located. The shell plumbing is identical; the archive formats differ.
  • Would bash's named-pipe fallback on a system without /dev/fd make unzip work?
    No. The fallback creates a real FIFO in a temporary directory, which gives the path a filesystem entry but not file semantics — a FIFO holds no contents, has no size, and cannot be seeked. `unzip` fails identically. Only materialising the bytes into an ordinary file changes anything.
  • Why does `sudo mytool <(generate_config)` fail with a missing-file error when the same line without sudo works?
    `sudo` closes inherited file descriptors above stderr before executing the target program, so the descriptor that `/dev/fd/63` names no longer exists by the time the tool opens it. The path is resolved by the *opening* process against its own table. Write the generated content to a temp file with restrictive permissions and pass that path instead.

saying these in an interview costs you the question

  • Blames the URL or curl rather than the pipe's semantics
  • Thinks /dev/fd/63 is a real file bash created
  • Believes the FIFO fallback makes the path seekable
  • Says unzip fails only because the download is incomplete
  • Expects the /dev/fd path to be meaningful over ssh

context