Why does an automatic merge need the common ancestor of the two lines of history?
answer
- Two texts show difference, not authorship
- Added on one side or removed on the other?
- A shared starting point attributes each change
- Three inputs, not two, hence three-way
- Most recent shared point, not the last release
basics
~20 sThe common ancestor tells the merge which side changed what. Two versions alone are ambiguous: a difference could be an addition on one side or a removal on the other. Comparing both to the shared start resolves it.
solid answer
~40 sTwo versions on their own show *that* they differ, not *who changed what*. A line present in one and absent in the other could mean one side added it or the other side removed it, and those two histories demand opposite results. The **common ancestor** — the most recent point both lines descend from — is the version both sides started from, so comparing ancestor-to-one-side and ancestor-to-the-other turns an ambiguous difference into an attributed change. Every case except *both sides changed this region differently* then has a determined answer. It also shapes resolution: reading the ancestor's version of a conflicted region tells you what each author was reacting to, and it is why the correct resolution is often text that appears on neither side.
code
pseudocode · 6 linesancestor line 12: reviewer_note = required
side A line 12: (line absent)
side B line 12: reviewer_note = required
two-way view the two sides differ at line 12 -> who is right? unknowable
three-way view ancestor had it, A removed it, B left it alone -> remove itgo deeper
Recall that merging looks at three versions, not two: where both sides started, and where each side ended up. Being able to say why a difference alone is ambiguous — added here or removed there — is enough at this level.
Expect to walk the attribution table row by row and explain why each row is decidable. Interviewers also ask which point is chosen as the base, so be ready to say most recent shared point, and why an older one would manufacture conflicts.
Demonstrate that you use the ancestor while resolving, not just while explaining. Talk about reading what each author was reacting to, producing merged text that belongs to neither side, and re-checking a resolution made partway through a replayed series.
Own the consequences at scale: histories where lines repeatedly cross can have several candidate ancestors and produce conflicts that surprise people. Be ready to say how you keep integration shapes simple enough that resolution stays a local, explainable decision.
## Two versions are not enough information Put two versions of a file side by side and you can see *that* they differ. You cannot see *who changed what*. A line present on one side and absent on the other has two equally plausible histories: one side added it, or the other side removed it. Those histories demand opposite results — keep the line, or drop it — and nothing in the two texts distinguishes them. The **common ancestor**, or merge base, is the most recent point in history that both lines descend from. It is the version both sides started from, so it converts an ambiguous difference into an attributed change. Comparing ancestor-to-one-side and ancestor-to-the-other, rather than side against side, is what makes automatic merging possible at all; that is why the technique is called a **three-way** merge. ## What the third input decides | At the ancestor | Side A | Side B | Attributed as | Result | |---|---|---|---|---| | line present | present | absent | B removed it | remove it | | line present | absent | present | A removed it | remove it | | line absent | present | absent | A added it | add it | | line present | changed | unchanged | A changed it | A's text | | line present | changed | changed differently | both changed it | conflict | Every row except the last is settled without asking anyone, and a two-way comparison could produce none of them with confidence. The ancestor is not a tiebreaker; it is the thing that makes the question decidable. ## Choosing the ancestor, and when there is more than one The base is the most recent point both tips descend from — not the branch point of whichever line is older, and not the last shared release. Picking anything further back would attribute changes that were already agreed as if they were new, and the merge would ask about regions nobody touched twice. When two lines have exchanged content before — each has already taken changes from the other — there can be more than one equally recent common ancestor. Implementations either pick one or build a synthetic base by first merging the candidates. The practical consequence is that history where lines repeatedly cross can produce conflicts that look surprising, because the version you are being compared against is not the point you think of as *where we split*. ## When the roles invert The labels *the side already in place* and *the incoming side* are assigned by which content is currently checked out, not by who wrote it. In a straightforward merge the content in place is the line you are standing on, so the in-place side is your team's work. When a series of recorded commits is instead **replayed** one at a time onto an updated line, that flips: the content in place is the other line plus whatever of your commits have already been replayed, and the incoming side is the single commit being applied right now. People routinely resolve a replay backwards for exactly this reason. Two further consequences are worth knowing. Each replayed commit is a separate resolution, so one stubborn region can require a decision several times over. And a decision made partway through applies to an intermediate state that never existed on either original line, which is why a replay resolution deserves a second look once the whole series has landed. ## What the ancestor cannot do Attribution is textual. The ancestor tells the algorithm which side edited a region; it says nothing about whether the two edits mean the same thing, contradict each other, or are both wrong. Two sides that rename the same concept differently conflict noisily. Two sides where one renames a concept and the other adds a use of the old name do not conflict at all, because those edits sit in different regions — the ancestor cannot rescue you there, and only building and exercising the merged result will. ## Resolution is not a menu of two options Because the ancestor attributes changes rather than ranking them, the correct resolution is frequently neither side verbatim. If one side changed a field's meaning and the other side added a caller relying on the old meaning, the merged text has to express both intents — content that exists on neither side. Reaching for a whole-side choice is the most common resolution mistake: it is fast, it merges clean, and it silently deletes half the intent. Two habits follow: 1. **Read the ancestor's version of the conflicted region before deciding** — it tells you what each author was reacting to. 2. **Treat the merged file as code nobody has ever run** — build it and exercise it before sharing it.
- When a series of commits is replayed one at a time onto an updated line, which content counts as the side already in place?The target line plus whatever of the replayed commits have already been applied. The incoming side is the single commit being applied at that moment, so the roles feel inverted compared with a merge. Because each commit is resolved separately, one region can demand a decision repeatedly, and each decision applies to an intermediate state that existed on neither original line.
- Can a resolution legitimately produce text that appeared on neither side?Frequently, and it is usually the sign of a good resolution. The ancestor attributes two intents without ranking them; when both are still wanted, the merged region has to express both. Picking one whole side is faster and produces a clean result, but it quietly discards work that somebody wrote on purpose.
- Why is the last shared release a bad choice of ancestor?Because everything agreed between that release and the real branch point would be re-attributed as new changes on both sides. Regions nobody touched twice would surface as conflicts, and resolving them risks reverting agreed work. The base has to be the most recent point both tips descend from.
Two colleagues mark up copies of the same printed letter. Holding just the two marked copies, you cannot tell who struck a sentence out and who left it alone; holding the original alongside them, every mark has an author and the merge becomes mechanical.
saying these in an interview costs you the question
- Thinks a merge compares only the two current versions
- Cannot explain what the common ancestor is for
- Assumes any older shared point works as the base
- Says the side with the newer timestamp should win
- Believes resolution must pick one side verbatim
- Never opens the ancestor version of a conflicted region