skip to content

In git blame, what do the -M and -C options change about the result?

level: middleimportance: should knowfreq 35%

answer

  1. Git records snapshots, not moves
  2. Refactorer becomes the apparent author
  3. One letter for within-file
  4. Repeat the letter to widen the search
  5. Add whitespace tolerance to help detection

basics

~20 s

Both stop a move from being reported as new authorship. git blame -M detects lines moved or copied within the same file, and -C additionally detects lines that came from other files changed in the same commit; repeating -C widens the search further.

solid answer

~50 s

Without them, moving code counts as deleting it in one place and writing it in another, so the person who reorganized a file becomes the apparent author of everything they moved. `git blame -M` makes Git look for the moved or copied lines elsewhere in the *same* file and attribute them to their original commit. `git blame -C` does that and also searches files that were modified in the same commit — the case where you extract a helper into a new module. Repeating the option widens the hunt: `-C -C` also inspects the files present in the commit that created the file, and `-C -C -C` looks at all files in every commit considered, which is slow but thorough. Both accept an optional numeric argument setting how many alphanumeric characters a run must have before it is considered a match. Pair them with `-w` so reindentation during the move does not defeat detection.

go deeper

for a junior

Recognize that a refactoring commit can hide real authorship and that git blame has options to look through it; the exact flag semantics are a level up.

for a middle

State precisely what each option searches: -M within the file, -C also files modified in the same commit, and repetition widening the scope. Mention combining with -w.

for a senior

Show a diagnostic routine — escalate -w, -M, -C, -C -C only as needed — and explain the cost and false-attribution tradeoffs that keep them off by default.

for a principal

Tie it to how the team commits: separating pure moves into their own commits, with honest messages, makes archaeology cheap for everyone and reduces reliance on expensive detection.

## The failure they fix Git stores snapshots, not moves. When you cut fifty lines out of `utils.py` and paste them into `parsers.py`, the commit records fifty deleted lines in one file and fifty added lines in the other. Plain `git blame parsers.py` then attributes every one of those lines to you and to the refactoring commit, erasing years of real authorship. The same happens when you merely reorder functions inside a single file. This matters because the whole point of blame is archaeology: you want the commit that explains *why the logic is what it is*, and a move commit explains nothing. ## -M: within the file `git blame -M <file>` tells blame that when it sees a block of added lines, it should first look for the same content elsewhere in the *same file* in the parent commit. If it finds it, the lines are attributed to whatever commit last touched them in their old position rather than to the move. The option accepts an optional threshold, written attached to the letter — `-M40` — giving the minimum number of alphanumeric characters a moved run must contain before Git will treat it as a move. Raising it suppresses coincidental matches on short, common lines; lowering it catches smaller moves at the cost of noise. ## -C: across files `git blame -C <file>` includes everything `-M` does and additionally looks for the lines in **other files that were modified in the same commit**. This is the extract-a-helper case: you created `parsers.py` and edited `utils.py` in one commit, so `utils.py` is in scope and the origin is found. Repetition widens the search: - **`-C`** — other files modified in the same commit. - **`-C -C`** — additionally, files present in the commit that *created* the file being blamed. This catches the common pattern where a file was created by wholesale copying from elsewhere and the source file was not touched in that commit. - **`-C -C -C`** — additionally, all files in every commit examined. Exhaustive and correspondingly slow on a large repository. Like `-M`, `-C` takes an optional numeric threshold. ## Why they are not the default Cost. Each level of detection makes blame examine more content per commit, and blame is already a history walk. On a large repository with deep history, `-C -C -C` on a big file can take a long time. Git's default keeps the common case fast, and you opt in when the answer looks wrong. There is also a correctness argument: aggressive copy detection can attribute a line to a distant file that merely happened to contain similar boilerplate. The thresholds exist to manage exactly this. ## Combining with whitespace tolerance Moves are rarely pure. Code extracted into a function usually gets reindented, so the moved lines are not byte-identical and detection can fail. Adding `-w` makes blame ignore whitespace when comparing, which dramatically improves hit rate on real refactors. `git blame -w -C -C <file>` is a reasonable "try harder" invocation when the plain result points at a refactoring commit. ## Reading the output When `-C` finds a line's origin in a different file, blame shows the **original path** alongside the commit for that line, so you can see where the code lived before. That path is often the most valuable part of the output: it tells you which module the logic used to belong to, which in turn tells you what it was originally for. ## Related but distinct machinery `git diff` has its own rename and copy detection (`-M`/`-C` there too, with `diff.renames` enabled by default in modern Git) which operates at whole-file granularity — it decides that file A became file B. Blame's `-M`/`-C` operate at *line* granularity within and across files. They solve related problems with similar spellings, and conflating them is a common mistake: whole-file rename detection does not help when a function moved between two files that both still exist. For reading a diff rather than a blame, `--color-moved` is the analogous readability tool, colouring moved blocks so a reviewer can see at a glance that a hunk is a relocation. ## Practical recipe When a blame result points at a commit whose message is a refactor, escalate in this order: add `-w`, then `-M`, then `-C`, then `-C -C`. Stop as soon as you get a commit that plausibly explains the logic, and read that commit with `git show`.

  • Why are -M and -C not enabled by default?
    Cost, mainly: each level makes blame inspect more content per commit on top of an already expensive history walk, and the widest form can be very slow on large repositories. Aggressive copy detection can also attribute lines to unrelated files containing similar boilerplate, which is why the thresholds exist.
  • What extra information appears in blame output when -C finds a line's origin in another file?
    The original path is shown alongside the commit for that line, so you can see where the code lived before the move. That path is often the most useful part of the result: it tells you which module the logic originally belonged to and therefore what problem it was written to solve.
  • Why does adding -w often make -C succeed where it failed alone?
    Because extracted code is usually reindented, so the moved lines are not byte-identical and detection misses them. -w makes the comparison ignore whitespace, so a block that was only shifted in indentation still matches its origin.
  • How does blame's -M differ from git diff's -M?
    They work at different granularity. In git diff, -M is whole-file rename detection: it decides that file A became file B. In git blame, -M finds individual moved or copied lines within a file, and -C extends that across files. Whole-file detection does not help when one function moved between two files that both still exist.

saying these in an interview costs you the question

  • Thinks Git records moves explicitly
  • Believes -M and -C are on by default
  • Confuses blame's -M with diff rename detection
  • Assumes -C searches the entire repository history
  • Gives up when blame points at a refactor commit

context