skip to content

A colleague runs `sort data.txt > data.txt` in bash and data.txt comes back empty. Why does the file lose its contents, and what is the correct way to filter a file in place?

level: middleimportance: must knowfreq 62%

answer

  1. who opens the file, and when
  2. expansion, redirection, then execution
  3. O_TRUNC fires at setup time
  4. the input is gone before sort starts
  5. temp file plus atomic rename

basics

~20 s

The shell performs redirections before it runs sort, so > data.txt truncates the file to zero bytes first and sort then reads an empty file. Filter into a temporary file and rename it over the original instead.

solid answer

~50 s

Redirection is done by the shell, not by the command. For a simple command bash expands the words, sets up every redirection, and only then executes the program — and `>` opens its target with `O_TRUNC`, so `data.txt` is emptied before `sort` is even started. `sort` opens the file it was given, finds zero bytes, and dutifully writes nothing. No tool can defend against this from the inside, because the damage is done before it runs. The portable fix is a temporary file plus a rename: `sort data.txt > data.tmp && mv data.tmp data.txt`, which is also atomic within one filesystem. Some tools have a safe in-place mode instead — GNU `sort -o data.txt data.txt` is explicitly documented to allow the same file, and `sed -i` writes a temp file and renames it for you.

go deeper

for a junior

Recognise the shape cmd file > file as destructive and reach for a temporary file plus mv. Being able to say the shell empties the file before the command starts is enough at this level.

for a middle

Walk through the order bash uses for a simple command — expand words, apply redirections, execute — and tie O_TRUNC to the empty file. Contrast > with >> and name one safe in-place option such as sort -o.

for a senior

Show the durability angle: && so a failed filter does not replace good data, the temp file on the same filesystem so the rename is atomic, and awareness that sed -i changes the inode and breaks hard links.

for a principal

Frame it as a data-loss class, not a trivia answer: mandate a write-temp-then-rename convention for anything that rewrites files in place, and decide where a shell filter stops being an acceptable tool for mutating production data at all.

## The shell does the redirection, not the command The single most useful mental model here is that `>` is not an argument, an option, or something the program interprets. When bash runs a simple command it works through a fixed order: it splits the line into words, performs expansions, **removes the redirections from the word list and applies them**, and only then executes the program with the remaining words as its arguments. `sort` receives exactly one argument — `data.txt` — and inherits a descriptor 1 that is already attached to a freshly emptied `data.txt`. Because `>` opens the target with the `O_TRUNC` flag, the truncation happens at *setup* time. By the time `sort` calls `open("data.txt")` to read its input, the file it opens is the same, now zero-byte, file. There is no race, no timing subtlety and no version where this works: `cmd file > file` reliably destroys `file`. ```bash printf 'b\na\n' > data.txt sort data.txt > data.txt # data.txt is now empty ``` The same trap wears other costumes: `grep -v DEBUG app.log > app.log`, `tr -d '\r' file > file`, `jq '.' conf.json > conf.json`. Any filter whose input and output name the same path loses the input. ## Why the tool cannot save you Candidates sometimes propose "sort should detect this". It mostly cannot, because by the time it starts, the previous contents are already gone — the shell truncated the file in a process that has since exec'd into `sort`. A tool *can* detect that its input and output paths refer to the same file when it is asked to open both itself, which is precisely why GNU `sort` offers `-o`: ```bash sort -o data.txt data.txt # documented as safe: sort reads the input fully first ``` Here `sort` owns the output path, so it can read the input completely before opening the output for writing. That guarantee comes from `sort`, not from the shell. ## The correct patterns **1. Temp file plus rename.** The workhorse, and the only one that works with arbitrary tools: ```bash sort data.txt > data.tmp && mv data.tmp data.txt ``` The `&&` matters: if the filter fails, you keep the original instead of replacing it with a half-written file. A rename within the same filesystem is atomic, so any concurrent reader sees either the old file or the new one, never a truncated one. Create the temporary next to the target — not in `/tmp` — so the rename stays on the same filesystem; a cross-device `mv` degrades into a copy plus delete and loses atomicity. (Generating that temp name safely, and deleting it when the script dies, is its own discipline.) **2. A tool's own in-place mode.** `sed -i`, `perl -i`, `sort -o`, `sponge` from moreutils. Note that `sed -i` is not magic either: it writes a temporary file and renames it, which is why it changes the file's inode and can break hard links. And `sed -i` needs an argument on BSD/macOS (`sed -i '' …`) where GNU takes none. **3. Read it all first.** `data=$(cat data.txt)` then write — fine for small files, hopeless for large ones, and it mangles trailing newlines. ## noclobber as a guardrail Bash can refuse to overwrite an existing file: ```bash set -o noclobber # same as set -C date > data.txt # bash: data.txt: cannot overwrite existing file date >| data.txt # >| overrides noclobber for this one redirection set +o noclobber ``` Under `noclobber`, `sort data.txt > data.txt` fails loudly instead of silently emptying the file — a genuine save. Know its limits: it only affects `>` (not `>>`, and not tools that open files themselves), it is a per-shell option that scripts must set for themselves, and it is off by default in every shell you will meet. Treat it as a seatbelt for interactive sessions, not as a substitute for the temp-file-and-rename pattern in a script. ## What an interviewer is really testing This question separates people who have memorised operators from people who know *when* each stage of command execution happens. Say the order out loud — expand, redirect, execute — and the empty file explains itself. It also generalises: the same ordering is why a redirection error aborts the command before it runs at all, and why `> file` creates the file even when the command turns out not to exist.

  • Why is `sort data.txt > data.tmp && mv data.tmp data.txt` safer than piping through a second copy of the file?
    Two reasons. The `&&` means a failing filter leaves the original untouched instead of replacing it with partial output, and `mv` within one filesystem is an atomic rename, so a concurrent reader sees either the complete old file or the complete new one. Keep the temp file beside the target, or the rename becomes a copy and loses that atomicity.
  • Does `set -o noclobber` prevent every accidental overwrite in a script?
    No. It only guards the `>` operator in the shell that has it set: `>>` still appends, `>|` deliberately overrides it, and any tool that opens an output file itself — `tee`, `sed -i`, `cp` — is unaffected. It is a useful interactive seatbelt, not a substitute for writing to a temp file and renaming.
  • If the redirection target cannot be opened, does the command still run?
    No. The shell applies redirections before executing the command, so a failure to open the target aborts that command with a non-zero status and the program is never started. That is also why `> /some/unwritable/path` reports the error with the shell's name in the message rather than the tool's.

saying these in an interview costs you the question

  • Claiming sort truncates its own output file
  • Thinking output is written only after the command exits
  • Believing >> would have fixed it correctly
  • Assuming the tool can detect same-file input
  • Using /tmp for the temp file then calling mv atomic

context