skip to content

How does Git detect renames during a merge, and where does that detection break down?

level: seniorimportance: should knowfreq 38%

answer

  1. Nothing about renames is stored
  2. It is inferred from content
  3. Similarity threshold and a cap
  4. Heavy rewrites and splits defeat it

basics

~20 s

Commits store no rename records, so Git infers renames at merge time by matching content: identical blobs are exact renames, and remaining files are paired by similarity. Detection fails when a file changed too much, was split, or when the rename limit is exceeded.

solid answer

~50 s

Git never records that a file was renamed — a commit stores whole trees, so `git mv` is just a delete plus an add. During a merge, the strategy reconstructs renames by comparing content: files whose blob hashes match are exact renames, and remaining added and deleted paths are paired if their similarity passes a threshold, by default around 50%. When detection succeeds, the other branch's edits to the old path are carried onto the new path instead of being lost. It breaks down when a file was renamed *and* rewritten below the threshold, when one file was split into several, when both sides rename the same file differently (a rename/rename conflict), or when the number of candidate paths exceeds `merge.renameLimit` / `diff.renameLimit`, where Git warns and skips inexact detection. You can tune it with `-X find-renames=<n>` or disable it with `-X no-renames`.

code

bash · 3 lines
bash
git merge -X find-renames=30 feature
git merge -X no-renames feature
git -c merge.renameLimit=5000 merge feature

go deeper

for a junior

Know that Git does not store renames: it guesses them by comparing file content when it shows a diff or performs a merge.

for a middle

Explain the two stages — exact matching by blob hash, then similarity pairing above a threshold — and name that the threshold and the candidate limit are configurable.

for a senior

Diagnose a merge where a moved file's edits vanished: check for the rename-limit warning, retry with a lower find-renames threshold, and enforce pure-move commits going forward.

for a principal

Own restructure planning: sequence large moves as standalone commits and set expectations for in-flight branches, because detection is a heuristic you can help or defeat.

## Git stores no renames This is the fact everything else follows from. A Git commit points at a tree, and a tree lists paths and blob ids. There is no "renamed from" field anywhere in the object model. `git mv` is a convenience that stages a deletion of one path and an addition of another; the resulting commit is indistinguishable from doing it by hand. Renames are therefore **inferred at read time** — by `git log --follow`, by `git diff`, and by the merge strategy — from the content itself. ## Why a merge cares Suppose one branch renames `parser.go` to `syntax/parser.go` while another branch fixes a bug inside `parser.go`. Without rename detection, the merge sees a file deleted on one side and modified on the other: a modify/delete conflict at best, a silently dropped bug fix at worst. With detection, Git recognises the rename and applies the other side's edits to the new path, producing the result a human would expect. ## How detection works Two stages: 1. **Exact renames.** A deleted path and an added path whose blob object ids are identical are the same content at a new location. This is cheap — a hash comparison — and catches pure moves. 2. **Inexact renames.** For the remaining unmatched additions and deletions, Git measures content similarity between candidates and pairs those above a threshold, by default around 50% similar. This step is quadratic in the number of candidates, which is why it is capped. Git also performs **directory rename detection**: when one side moves a whole directory and the other side adds new files to the old directory, the new files can be moved into the renamed directory. Where this is ambiguous — for example the directory's files went to two different destinations — Git reports a conflict rather than guessing. ## The knobs - `-X find-renames=<n>` sets the similarity threshold for the merge, for example `-X find-renames=30` to match more aggressively at the cost of false pairings. - `-X no-renames` turns rename detection off for the merge. - `merge.renames` is the config equivalent for enabling or disabling merge rename detection. - `merge.renameLimit` and `diff.renameLimit` cap how many candidate paths inexact detection will consider. Exceeding the cap makes Git emit a warning that inexact rename detection was skipped, and the merge proceeds as if there were no renames. ## Where it breaks down - **Rename plus heavy rewrite.** Move a file and rewrite most of its body in the same commit and similarity drops below the threshold, so the pair is not detected. The other side's edits then have nowhere to go. - **Splits and merges of files.** Rename detection pairs one deletion with one addition. A file split into three modules, or three files merged into one, has no such pairing to find. - **Divergent renames.** Both sides rename the same file to different paths — a rename/rename conflict — and Git cannot choose for you. - **Rename versus delete.** One side renames, the other deletes: the conflict is real and requires a human decision about intent. - **Limit exceeded.** Large restructures produce thousands of candidates; when the rename limit is hit, detection is skipped wholesale and the merge degrades into add/delete pairs, often with a storm of conflicts. - **Generated or boilerplate files.** Highly similar files can be paired with the *wrong* partner, which is the false-positive side of the same heuristic. ## Practical mitigations The highest-value habit is **structural**: perform large moves in a commit that does nothing else. A pure-move commit is trivially detectable — the blobs are identical, so exact detection catches everything without touching the similarity stage or the rename limit. Mixing a move with edits is what pushes similarity down. When a merge has already gone wrong, the diagnostic path is to check whether Git warned about the rename limit, retry with a raised limit or a lower `find-renames` threshold, and inspect the conflicting paths to see whether they are actually a moved-and-edited pair. ## What an interviewer is testing The key insight is that rename tracking is *inference, not record-keeping*. Candidates who say "Git tracks renames with `git mv`" have the model wrong, and that wrong model makes every downstream failure inexplicable. Once you state the inference model, the failure cases follow naturally: anything that lowers content similarity, or makes the pairing one-to-many, or makes the search too expensive, defeats a heuristic that was never a guarantee in the first place.

  • Does git mv record the rename in the commit?
    No. `git mv` stages a delete of the old path and an add of the new one; the resulting commit is identical to doing both by hand. Any rename you see later in a diff, log or merge is reconstructed from content similarity at the time you look, not read from stored metadata.
  • How can you make a merge more aggressive about finding renames?
    Lower the similarity threshold with `-X find-renames=<n>`, and raise `merge.renameLimit` or `diff.renameLimit` if Git warned that inexact detection was skipped. A lower threshold increases false pairings, so treat it as a diagnostic tool for a specific difficult merge rather than a permanent setting.
  • What is the best way to keep a large directory restructure mergeable?
    Do the move in a commit that changes nothing else. Pure moves leave blob ids unchanged, so exact rename detection matches them without relying on the similarity stage or the rename limit. Edits mixed into the same commit lower similarity and are exactly what pushes a pair below the detection threshold.
  • What happens when both branches rename the same file to different paths?
    Git reports a rename/rename conflict. It has detected both renames but cannot decide which destination is intended, so it leaves the situation for a human. The resolution is a judgment call about which structure the project wants, not something a strategy option can settle correctly.

saying these in an interview costs you the question

  • Thinks Git stores rename metadata in commits
  • Believes git mv records the rename permanently
  • Assumes rename detection never fails
  • Thinks the similarity threshold is fixed and unconfigurable
  • Blames the merge tool when a moved file's edits are lost

context