skip to content

Branching and Merging

A branch is just a moving pointer, which is why fast-forward and three-way merges behave the way they do. Expect the merge-versus-rebase question and a request to explain what detached HEAD means and how you got there.

on this pageshow

questions

5

In a version-control system, what does a branch actually store, and why is creating one nearly free?

level: juniorimportance: must knowfreq 84%

answer

  1. It is far smaller than you assume
  2. A name, not a container
  3. Holds exactly one commit identifier
  4. It advances when you record work
  5. Ancestry comes from walking parent links

basics

~20 s

A branch stores nothing but a reference to one commit — a name that advances as you record work. Creating one writes a single tiny reference and copies no file content, so it is effectively instant.

solid answer

~50 s

A branch is a **named pointer into the commit graph**, not a copy of the project. The name holds the identifier of exactly one commit — its tip — and everything earlier is reachable by following parent links backwards from there. When you record a new commit while that name is current, the new commit's parent is the old tip and the name advances onto the new commit; nothing else in the repository moves. Creating a branch therefore writes one small reference and touches no file content, which is why it costs the same in a huge project as in a tiny one. Copying a working directory instead duplicates every file, and the two copies then share no recorded ancestry, so the system cannot reason about what they have in common. Cheap naming is what makes short-lived lines of work practical.

go deeper

for a junior

Be ready to say in one sentence that a branch is a name holding one commit identifier and that creating it copies no files. Interviewers use this as a screening check that you have a mental model at all.

for a middle

Explain the mechanics out loud: the new commit records the old tip as its parent, the name is updated to the new commit, and nothing else changes. Be able to say why the stored reference is the same size in any project.

for a senior

Show what the model buys in practice — reasoning about two lines by their common ancestor rather than by comparing folders, several names addressing one commit, and reclaiming an abandoned line by removing a reference.

for a principal

Own the tradeoff the cheapness hides: creation is constant-cost but integration cost grows with divergence, so the lever a lead actually controls is how long a line of work is allowed to live, not how many exist.

## What a repository actually stores A repository keeps its history as a **graph of commits**. Each commit is an immutable record holding a snapshot of the whole project at one moment plus the identifier of the commit or commits it came from. Because every commit names its parent, the entire history behind any point can be recovered by walking backwards along those parent links. Notice what is *not* in that description: there is no object in the graph called a branch. The graph is just commits pointing at their ancestors. A **branch is a name that holds one commit identifier** — the *tip* of a line of work. That is the whole of it. The stored reference is a few dozen bytes whether the project has forty files or four hundred thousand, because it stores an identifier and not content. ## Why creating one is nearly free The cost of an operation is the cost of the bytes it must write. Creating a branch writes one reference. It does not read the working tree, does not compress anything, and does not duplicate a single file. | | Copying the project directory | Creating a branch | |---|---|---| | Bytes written | every file, again | one small reference | | Time | grows with project size | effectively constant | | Shared ancestry | none recorded | full — same graph | | System can compare them | only file by file | by common ancestor | | Cleanup | delete a large tree | remove one name | The last two rows matter more than the speed. Two copied folders are two unrelated piles of files; the system knows nothing about how they relate. Two branch names point into **one shared graph**, so the system can ask a question a file comparison cannot answer: where did these two lines last agree? ## What happens when you record a commit Suppose a name points at commit `C` and that name is the current one. You record new work. Three things happen, in this order: 1. A new commit `D` is written, holding the new snapshot and recording `C` as its parent. 2. The branch name is updated to hold `D`'s identifier instead of `C`'s. 3. Nothing else changes. `C` is untouched and still reachable — it is `D`'s parent. This is why the name is described as *moving*. It is not the history that moves; the history only ever grows. The name slides forward along the newest commit, and everything behind it stays exactly as it was recorded. ## Consequences worth being able to state - **Many names can point at the same commit.** Creating a second branch at your current position costs another few dozen bytes and produces two names for one place. Nothing is duplicated. - **Deleting a branch name deletes a name.** The commits it pointed at are still stored. They may become unreachable if no other name or commit leads to them, but the deletion itself removes only the reference. - **A branch has no owner and no contents of its own.** Asking "which commits are on this branch" really means "which commits are reachable from this name", and a commit can be reachable from several names at once. That is why the same commit legitimately appears on more than one line of work. - **The name says nothing about time.** A branch name pointing at an old commit is not stale data; it is an accurate record that this line of work last advanced then. - **Cheap to create is not cheap to integrate.** The creation cost is constant, but the cost of bringing a line back grows with how far the two lines have diverged, because that is real content the system must reconcile. ## The misreadings to avoid The most common wrong model is that a branch is a *place where changes live* — a container you put work into, which is then merged back by pouring the container out. Under that model it is impossible to explain why creating a branch is instant, why deleting one is instant, or why two names can address the same commit. The pointer model explains all three without special cases. A second wrong model is that the copy is deferred and happens on first change. Nothing is deferred, because nothing was ever going to be copied. Content is written when you record a commit, and it is written once, keyed by what it contains — two branches holding an identical file hold the identifier of the same stored content, not two copies of it. The practical payoff is the point of the whole design: when starting a line of work costs nothing and abandoning one costs nothing, teams start them freely, and the interesting engineering question moves from "can we afford a branch" to "how long should any line be allowed to live before it is brought back".

  • If a branch stores only one identifier, how does the system know which commits are on that line of work?
    It does not store a list. It starts at the commit the name holds and walks parent links backwards, so "on this line" means "reachable from this name". Because a commit can be reached from several names, the same commit can legitimately belong to more than one line at once, with nothing duplicated.
  • What actually disappears when you delete a branch name?
    Only the reference. The commits it pointed at remain in the object store exactly as recorded. If no other name and no other commit leads to them they become unreachable and are eventually reclaimed, but that is a later storage decision, not part of the deletion. Deleting a name is not deleting history.
  • Two branch names point at the same commit. Is anything duplicated?
    No. Both names hold the same identifier and both walk into the same shared graph. The only extra cost is the second reference itself. This is also why comparing the two lines reports nothing to reconcile: they have not diverged, so their most recent common ancestor is the commit they both name.

A branch name is a bookmark in a book that is only ever appended to: moving the bookmark forward as new pages are written costs nothing, and adding a second bookmark on the same page does not duplicate the book.

saying these in an interview costs you the question

  • Says a branch copies the whole project into a new folder
  • Thinks branch creation gets slower as the repository grows
  • Believes a branch stores the changes made on it
  • Cannot say what the branch name points at
  • Claims deleting a branch name erases those commits from storage
  • Says two names on one commit means the work exists twice
open as a page

Why can't a version-control merge be decided by comparing only the two branch tips?

level: middleimportance: must knowfreq 67%

basics

~20 s

Two tips reveal that they differ but not who changed what. The system also reads the merge base — the most recent commit both lines share — so a change only one side made can be taken automatically instead of being questioned.

open as a page

When can integrating a version-control branch just move a pointer instead of creating a two-parent commit?

level: middleimportance: should knowfreq 53%

basics

~20 s

Only when the target has recorded nothing since the two lines parted — its tip is an ancestor of the source's tip. Then there is nothing to combine and the name simply advances. If both lines moved, a commit with two parents is required.

open as a page

In version control, how do you choose between a merge commit that records the true graph and replaying commits onto a straight line?

level: seniorimportance: should knowfreq 63%

basics

~20 s

Choose by what the history must answer later. A merge commit preserves what actually happened and keeps existing commit identifiers stable; replaying produces a linear story that never happened and writes new commits, so it is safe only on a line nobody else has copied.

open as a page

In version control, what does it mean to sit on a commit that no branch name points to?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

Your position is a specific commit rather than a moving name. Inspecting is fine, but anything you record there gets no name, so once you move away nothing refers to it and the work becomes unreachable and eventually reclaimed.

open as a page