When is a custom .gitattributes driver the wrong way to solve a Git repository problem?
answer
- One half is committed, the other is not
- Config is never transferred by a clone
- There is a security reason for that split
- Built-ins need no per-machine setup
- Ask whether absence fails loudly or silently
basics
~20 sCustom diff, merge and filter drivers are named in a committed file but defined in per-machine Git config, so behaviour silently differs by machine. When correctness depends on the driver, prefer built-in attributes or remove the artifact from the repository.
solid answer
~50 sThe asymmetry decides it. `.gitattributes` is committed, so the *assignment* reaches everyone, while `diff.<name>.textconv`, `merge.<name>.driver` and `filter.<name>.clean` live in config, which is never transferred by clone and which Git deliberately will not accept from a repository, because that would let checked-out content execute commands. So a custom driver produces a split brain: contributors who ran the setup step get one behaviour, everyone else and every automated checkout get the default, usually with no warning. That is acceptable for conveniences such as nicer diffs, and unacceptable when correctness depends on it, for instance a filter that keeps large payloads or secrets out of commits, or a merge driver that prevents a corrupt result. My rule is: built-in attributes that need no config (`-diff`, `binary`, `merge=union`, `export-ignore`) can be committed freely; anything requiring config needs a bootstrap step, a `required` flag where one exists, and a check that failure is loud rather than silent.
code
gitattributes · 3 linesdocs/**/*.pdf diff=pdf
assets/** binary
CHANGELOG.md merge=uniongo deeper
Know that writing an attribute line is only half the setup: the driver it names must also exist in each person's Git configuration, so a colleague may see different behaviour from the same repository.
Explain which behaviours are built in and need no configuration, and which require a definition in config, and be able to predict what a teammate without that definition actually observes.
Show that you plan for the population without the driver, especially automated checkouts, and that you make absence either harmless or loud rather than silently wrong.
Own the decision framework: question whether the artifact belongs in the repository, prefer built-ins that degrade safely, treat any custom driver as a permanent dependency on every machine and image, and name the security reason Git refuses to take executable configuration from a clone.
## The structural problem Attributes and drivers live on opposite sides of a trust boundary. A `.gitattributes` line is data in the repository and travels with a clone. A driver is a command line in a config file, on the user's machine, and Git will not read executable configuration out of the checked-out tree. That separation is a security property, not an oversight: if cloning a repository could install commands that later run during `git add` or `git checkout`, cloning would be equivalent to executing untrusted code. The cost of that safety is that the assignment and the behaviour ship separately, and only one of them is versioned. ## Failure modes to reason about - **Silent divergence.** Someone without the driver gets the default. A merge driver falls back to the ordinary text merge; a textconv assignment produces a plain binary-differ notice. Nothing warns them that the repository expected something else, so results differ by machine without anyone noticing. - **Automated checkouts.** Build and release environments start from a clean image and have none of the developer setup, so they are the population most likely to run without the driver, and the least likely to notice. - **Hard failure at the wrong moment.** Setting a filter's `required` flag converts silence into an abort, which is usually correct, but it also means a contributor without the filter cannot check the repository out at all. That tradeoff has to be a decision, not an accident. - **Performance.** Filters run on every add and checkout of a matching path. On a broad pattern in a large tree, the per-file process cost is real; the long-running `process` form exists precisely for this, and choosing it is part of the design, not a later optimisation. - **Maintenance.** A driver is a dependency: a program that must exist, on the path, in a compatible version, on every developer machine and every build image, effectively forever. ## A decision order 1. **Should the artifact be in the repository at all?** A generated file that conflicts constantly, or a binary blob that bloats clones, is often better produced at build time. This removes the problem rather than automating around it. 2. **Can a built-in do it?** `-diff` for unreviewable text, `binary` for content that must never be merged, `merge=union` for append-only lists, `export-ignore` for archive contents: these are behaviours Git implements itself, so they need no per-machine setup, degrade to nothing surprising, and are safe to commit unilaterally. 3. **Is the custom behaviour a convenience or a correctness requirement?** Convenience, such as a prettier diff for a document format, can live in per-user config and be documented as optional. Correctness cannot: if a wrong result is possible when the driver is absent, either make its absence loud or choose a different mechanism. 4. **If a custom driver is still right**, then make setup mechanical: a checked-in script or a documented `git config` invocation, applied in developer onboarding and in the build image, with the `required` flag where the mechanism supports it, and a note in the repository explaining what the attribute expects. ## Alternatives that avoid the coupling Most problems people solve with custom drivers have a structural alternative. Constant conflicts in a generated index are better solved by splitting it into per-change fragment files that are combined at release time. Unreviewable generated code is better solved by not committing it, or by marking it `-diff` so review output stays readable without adding any dependency. Content that should never enter history is better handled by refusing it at the boundary, since a filter that is merely absent will not stop anything. ## What a strong answer sounds like The interviewer is checking whether you notice the versioning asymmetry unprompted, and whether you distinguish mechanisms that degrade safely from those that degrade silently into wrong results. Naming the security reason Git keeps executable configuration out of the repository, and then applying a preference for built-ins with an explicit escalation path, is the shape of the answer that lands.
- Why does Git refuse to read driver definitions from a file inside the repository?Because those definitions are command lines that Git executes during ordinary operations such as add, checkout and merge. If a clone could supply them, cloning an untrusted repository would be equivalent to running its code. Keeping executable configuration in per-user config files, which the user installs deliberately, is what stops a repository from choosing what runs on your machine.
- Which attribute-based behaviours are safe to commit without any per-machine setup?The ones Git implements itself: unsetting diff for unreviewable text, the binary macro, the built-in text, binary and union merge drivers, export-ignore and export-subst for archives, and the end-of-line attributes. They take effect for every clone and every automated checkout with no configuration, which is exactly why they should be the first thing you reach for.
saying these in an interview costs you the question
- Assumes committing the attribute is enough for the team
- Wants Git to read driver commands from the repository
- Ignores that build environments have no developer setup
- Treats a required filter's checkout failure as a free win
- Reaches for a custom driver before questioning the artifact