skip to content

In Git, what do clean and smudge filters do, and when does each one run?

level: seniorimportance: should knowfreq 32%

answer

  1. Two crossings of the working-tree boundary
  2. One direction stores, the other restores
  3. Commands are piped content on standard input
  4. A config flag decides whether failure is fatal
  5. This is the mechanism behind large-file storage

basics

~20 s

A clean filter transforms content on its way into the repository, when a file is staged; a smudge filter transforms it on the way out, at checkout. Both are named by a filter attribute and defined by commands in Git config.

solid answer

~50 s

Assign `filter=<name>` to a path in `.gitattributes`, then define the pair in config: `filter.<name>.clean` runs when content moves from the working tree into the object store, so the blob Git stores is the *cleaned* form, and `filter.<name>.smudge` runs on checkout, converting the stored blob back into the working-tree form. Both commands read the content on standard input and write the transformed content to standard output, with `%f` available as the path. By default a failing filter is tolerated and the content passes through unchanged; set `filter.<name>.required = true` to make failure abort the operation instead, which is the safe choice whenever the stored form is not usable as-is. `filter.<name>.process` selects a long-running protocol so one process handles many files rather than forking per file. This clean/smudge pair is the mechanism Git LFS is built on: the clean step replaces content with a small stored stand-in.

code

gitattributes · 2 lines
gitattributes
*.ipynb filter=stripout
*.psd filter=bigfile

go deeper

for a junior

Know that Git can transform file content automatically on the way into and out of the repository, and that this is configured per path through the attributes file. The names clean and smudge are worth memorising with their directions.

for a middle

Explain that clean runs at staging and produces what is stored, smudge runs at checkout and produces what lands on disk, that content is piped through the commands, and that the definitions live in config while the assignment is committed.

for a senior

Demonstrate the operational care: make the pair a deterministic inverse or live with phantom modifications, set the required flag when the stored form is a stand-in, use the long-running process form on broad patterns, and remember existing files are not reprocessed.

for a principal

Own the rollout risk: a filter is unversioned per-machine configuration that every contributor and every automated checkout must install, so decide deliberately whether the problem justifies that coupling or whether the content should simply not be committed.

## The two directions Content crosses the boundary between working tree and object store twice, and a filter can intercept each crossing. - **clean** runs on the way **in**, when Git turns a working-tree file into a blob, most visibly during `git add`. Whatever the clean command writes to standard output is what Git stores and hashes. - **smudge** runs on the way **out**, when Git writes a blob into the working tree during checkout. Its output is what lands on disk. The attribute assignment `*.dat filter=myfilter` says which paths participate; the commands live in config under `filter.myfilter.clean` and `filter.myfilter.smudge`. Content is piped on standard input and the result read from standard output, with `%f` expanding to the path if the command needs it. ## The round-trip requirement The pair should be an inverse: applying smudge and then clean should return the original blob. If it does not, Git sees the working-tree file as modified immediately after a fresh checkout, and `git status` reports phantom changes that never go away. Non-determinism causes the same symptom, so a clean command that embeds a timestamp or a hostname is a trap. ## Failure handling By default, if a filter command fails or is missing, Git treats the content as unfiltered and continues. That is friendly for optional transformations and dangerous for anything else: a missing clean filter means the raw content gets committed, which may be exactly the large or secret payload the filter existed to keep out. Setting `filter.<name>.required = true` turns a filter failure into a hard error that aborts the checkout or the add, which is the right default whenever the stored form is a stand-in rather than the real content. ## Performance The simple form forks the filter command once per file, which is fine for a handful of paths and painful on a large checkout. `filter.<name>.process` names a single long-running command that speaks a packet protocol and handles many files over one process lifetime. Reach for it when the filter is on a broad pattern. ## What it is used for The canonical use is storing a stand-in instead of the real payload: the clean filter replaces large content with a small descriptor that is what actually gets committed, and the smudge filter fetches the real content back at checkout. That is precisely the mechanism Git LFS is built on. Other uses include normalising generated content so it does not churn, and stripping volatile metadata such as notebook execution counts before storage. Keyword expansion, in the style of `$Id$`, is the classic cautionary example. Git's own documentation warns against it: it makes the stored content depend on where it was checked out, which fights the content-addressed model and confuses diffs. ## Two operational gotchas First, filters run when content crosses the boundary, not when you change the attributes file. Adding or changing a filter does nothing to files already in the index and the working tree; they keep their existing stored form until they pass through Git again. Making the change take effect means deliberately re-processing the affected paths. Second, and most important architecturally: the attribute is committed but the filter command is config, and config is per machine and never transferred by clone. A contributor who has not installed and configured the filter silently gets different behaviour, or, with `required` set, a checkout that fails outright. Neither is discoverable from the repository alone, so a repository that depends on a filter needs a documented setup step or a bootstrap script, and everyone including CI must run it. ## Debugging Confirm the attribute with `git check-attr filter -- <path>`, then confirm the config with `git config --get filter.<name>.clean`. Comparing what is stored against what is on disk is the fastest way to see which half of the pair is misbehaving: inspect the stored blob directly rather than trusting the working-tree copy, because the working-tree copy is by definition the smudged form.

  • What goes wrong if a Git clean filter is not deterministic?
    Git stores whatever the clean command produced, then compares later runs against it. If the command embeds anything variable, such as a timestamp, the cleaned form differs on the next comparison and `git status` reports the file as modified immediately after a fresh checkout. The change cannot be resolved by committing, because the next comparison differs again. Clean must be a pure function of the content.
  • Why does filter.<name>.required = true matter for a filter that replaces content with a stand-in?
    Without it, a missing or failing filter is tolerated and content passes through unchanged, so the real payload gets committed rather than the stand-in, or the stand-in is checked out as literal file content. Setting required turns that into a hard failure of the add or checkout, which is the discoverable outcome. It also means a contributor without the filter installed is told immediately rather than producing bad commits.

Clean is the vacuum-packing machine at the warehouse door and smudge is the person who unpacks it on the way out. If the two disagree about the packing, every shipment looks tampered with.

saying these in an interview costs you the question

  • Swaps the directions of clean and smudge
  • Assumes the filter ships with the repository
  • Believes a failing filter always aborts the operation
  • Expects changing attributes to reprocess existing files
  • Writes a non-deterministic clean command

context