skip to content

Version Control

Version control as a concept and as daily practice: commits and history, branching and merging, the distributed model, team workflows, and untangling conflicts. Every engineering interview assumes fluency here.

on this pageshow

explore

questions

25

In a version-control system, what does a branch actually store, and why is creating one nearly free?

level: juniorimportance: must knowfreq 84%

answer

  1. It is far smaller than you assume
  2. A name, not a container
  3. Holds exactly one commit identifier
  4. It advances when you record work
  5. Ancestry comes from walking parent links

basics

~20 s

A branch stores nothing but a reference to one commit — a name that advances as you record work. Creating one writes a single tiny reference and copies no file content, so it is effectively instant.

solid answer

~50 s

A branch is a **named pointer into the commit graph**, not a copy of the project. The name holds the identifier of exactly one commit — its tip — and everything earlier is reachable by following parent links backwards from there. When you record a new commit while that name is current, the new commit's parent is the old tip and the name advances onto the new commit; nothing else in the repository moves. Creating a branch therefore writes one small reference and touches no file content, which is why it costs the same in a huge project as in a tiny one. Copying a working directory instead duplicates every file, and the two copies then share no recorded ancestry, so the system cannot reason about what they have in common. Cheap naming is what makes short-lived lines of work practical.

go deeper

for a junior

Be ready to say in one sentence that a branch is a name holding one commit identifier and that creating it copies no files. Interviewers use this as a screening check that you have a mental model at all.

for a middle

Explain the mechanics out loud: the new commit records the old tip as its parent, the name is updated to the new commit, and nothing else changes. Be able to say why the stored reference is the same size in any project.

for a senior

Show what the model buys in practice — reasoning about two lines by their common ancestor rather than by comparing folders, several names addressing one commit, and reclaiming an abandoned line by removing a reference.

for a principal

Own the tradeoff the cheapness hides: creation is constant-cost but integration cost grows with divergence, so the lever a lead actually controls is how long a line of work is allowed to live, not how many exist.

## What a repository actually stores A repository keeps its history as a **graph of commits**. Each commit is an immutable record holding a snapshot of the whole project at one moment plus the identifier of the commit or commits it came from. Because every commit names its parent, the entire history behind any point can be recovered by walking backwards along those parent links. Notice what is *not* in that description: there is no object in the graph called a branch. The graph is just commits pointing at their ancestors. A **branch is a name that holds one commit identifier** — the *tip* of a line of work. That is the whole of it. The stored reference is a few dozen bytes whether the project has forty files or four hundred thousand, because it stores an identifier and not content. ## Why creating one is nearly free The cost of an operation is the cost of the bytes it must write. Creating a branch writes one reference. It does not read the working tree, does not compress anything, and does not duplicate a single file. | | Copying the project directory | Creating a branch | |---|---|---| | Bytes written | every file, again | one small reference | | Time | grows with project size | effectively constant | | Shared ancestry | none recorded | full — same graph | | System can compare them | only file by file | by common ancestor | | Cleanup | delete a large tree | remove one name | The last two rows matter more than the speed. Two copied folders are two unrelated piles of files; the system knows nothing about how they relate. Two branch names point into **one shared graph**, so the system can ask a question a file comparison cannot answer: where did these two lines last agree? ## What happens when you record a commit Suppose a name points at commit `C` and that name is the current one. You record new work. Three things happen, in this order: 1. A new commit `D` is written, holding the new snapshot and recording `C` as its parent. 2. The branch name is updated to hold `D`'s identifier instead of `C`'s. 3. Nothing else changes. `C` is untouched and still reachable — it is `D`'s parent. This is why the name is described as *moving*. It is not the history that moves; the history only ever grows. The name slides forward along the newest commit, and everything behind it stays exactly as it was recorded. ## Consequences worth being able to state - **Many names can point at the same commit.** Creating a second branch at your current position costs another few dozen bytes and produces two names for one place. Nothing is duplicated. - **Deleting a branch name deletes a name.** The commits it pointed at are still stored. They may become unreachable if no other name or commit leads to them, but the deletion itself removes only the reference. - **A branch has no owner and no contents of its own.** Asking "which commits are on this branch" really means "which commits are reachable from this name", and a commit can be reachable from several names at once. That is why the same commit legitimately appears on more than one line of work. - **The name says nothing about time.** A branch name pointing at an old commit is not stale data; it is an accurate record that this line of work last advanced then. - **Cheap to create is not cheap to integrate.** The creation cost is constant, but the cost of bringing a line back grows with how far the two lines have diverged, because that is real content the system must reconcile. ## The misreadings to avoid The most common wrong model is that a branch is a *place where changes live* — a container you put work into, which is then merged back by pouring the container out. Under that model it is impossible to explain why creating a branch is instant, why deleting one is instant, or why two names can address the same commit. The pointer model explains all three without special cases. A second wrong model is that the copy is deferred and happens on first change. Nothing is deferred, because nothing was ever going to be copied. Content is written when you record a commit, and it is written once, keyed by what it contains — two branches holding an identical file hold the identifier of the same stored content, not two copies of it. The practical payoff is the point of the whole design: when starting a line of work costs nothing and abandoning one costs nothing, teams start them freely, and the interesting engineering question moves from "can we afford a branch" to "how long should any line be allowed to live before it is brought back".

  • If a branch stores only one identifier, how does the system know which commits are on that line of work?
    It does not store a list. It starts at the commit the name holds and walks parent links backwards, so "on this line" means "reachable from this name". Because a commit can be reached from several names, the same commit can legitimately belong to more than one line at once, with nothing duplicated.
  • What actually disappears when you delete a branch name?
    Only the reference. The commits it pointed at remain in the object store exactly as recorded. If no other name and no other commit leads to them they become unreachable and are eventually reclaimed, but that is a later storage decision, not part of the deletion. Deleting a name is not deleting history.
  • Two branch names point at the same commit. Is anything duplicated?
    No. Both names hold the same identifier and both walk into the same shared graph. The only extra cost is the second reference itself. This is also why comparing the two lines reports nothing to reconcile: they have not diverged, so their most recent common ancestor is the commit they both name.

A branch name is a bookmark in a book that is only ever appended to: moving the bookmark forward as new pages are written costs nothing, and adding a second bookmark on the same page does not duplicate the book.

saying these in an interview costs you the question

  • Says a branch copies the whole project into a new folder
  • Thinks branch creation gets slower as the repository grows
  • Believes a branch stores the changes made on it
  • Cannot say what the branch name points at
  • Claims deleting a branch name erases those commits from storage
  • Says two names on one commit means the work exists twice
open as a page

Why does a version-control system keep a staging area separate from the working copy?

level: juniorimportance: must knowfreq 74%

basics

~20 s

The staging area holds an explicit selection of what the next commit will contain, so you can record one logical change out of a messy working copy instead of committing everything you happen to have edited.

open as a page

What makes a version-control merge report a conflict instead of merging automatically?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A merge reports a conflict when both lines of history changed the same region of a file since their common ancestor. Changes in different regions combine automatically; two different changes to one region leave the algorithm no basis to choose.

open as a page

Why is every clone in a distributed version-control system a full repository, and what does that buy you?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A clone copies the entire history and every stored object, not just the newest files, so the copy is a working repository on its own — you can record changes, read history and switch lines of development with no server.

open as a page

Why is a version-control tag treated as a permanent name for one commit, while a branch keeps moving?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A tag names one exact commit in history and is meant never to move again; a branch is a pointer that advances to each new commit on its line. Re-pointing a released tag silently changes what consumers rebuild.

open as a page

Why require a change review before a version-control change lands on the shared mainline?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A change review is the last cheap checkpoint before one person's work becomes everyone's. It catches defects a machine has no model of, spreads knowledge of the code, and records why the change was accepted.

open as a page

Why can't a version-control merge be decided by comparing only the two branch tips?

level: middleimportance: must knowfreq 67%

basics

~20 s

Two tips reveal that they differ but not who changed what. The system also reads the merge base — the most recent commit both lines share — so a change only one side made can be taken automatically instead of being questioned.

open as a page

Why is a version-control commit a full snapshot rather than a stored diff?

level: middleimportance: must knowfreq 79%

basics

~20 s

A commit records the complete state of the file set at one moment, plus a pointer to the commit it was built on and who made it and why. Differences between two commits are computed on demand, never stored.

open as a page

When two copies of a distributed version-control repository synchronise, what actually moves between them?

level: middleimportance: must knowfreq 68%

basics

~20 s

Synchronising transfers the history objects one side lacks and then updates pointers. Retrieving from a shared copy moves your record of where their lines of development sit; your own line and your working files change only when you integrate.

open as a page

In semantic versioning, what does each of the three number positions promise a consumer?

level: middleimportance: must knowfreq 74%

basics

~20 s

Semantic versioning reads a version as major.minor.patch: a major bump warns of an incompatible change, a minor bump adds capability without breaking callers, and a patch bump is a backwards-compatible fix. The number is a compatibility promise, not marketing.

open as a page

What single axis separates branching models, and why does it decide everything else?

level: middleimportance: must knowfreq 76%

basics

~20 s

Branching models differ mainly in how long a line of development lives away from the shared mainline before integrating. Everything else follows from that, because the cost of divergence grows faster than the line's age.

open as a page

When can integrating a version-control branch just move a pointer instead of creating a two-parent commit?

level: middleimportance: should knowfreq 53%

basics

~20 s

Only when the target has recorded nothing since the two lines parted — its tip is an ancestor of the source's tip. Then there is nothing to combine and the name simply advances. If both lines moved, a commit with two parents is required.

open as a page

Why is a version-control commit's identifier derived from its own contents?

level: middleimportance: should knowfreq 46%

basics

~20 s

The identifier is a hash over the commit's snapshot, parent pointer, authorship metadata and message, so changing anything produces a different identifier. Two copies of history that share an identifier therefore share everything behind it.

open as a page

Why does an automatic merge need the common ancestor of the two lines of history?

level: middleimportance: should knowfreq 55%

basics

~20 s

The common ancestor tells the merge which side changed what. Two versions alone are ambiguous: a difference could be an addition on one side or a removal on the other. Comparing both to the shared start resolves it.

open as a page

How does a centralized version-control model differ from a distributed one, and what does centralization still do better?

level: middleimportance: should knowfreq 58%

basics

~20 s

A centralized model keeps the one history on a server and hands each developer a working copy of a single revision; a distributed model gives every copy the whole history. Centralization still wins at exclusive locking and a single global ordering.

open as a page

In version control, how do you choose between a merge commit that records the true graph and replaying commits onto a straight line?

level: seniorimportance: should knowfreq 63%

basics

~20 s

Choose by what the history must answer later. A merge commit preserves what actually happened and keeps existing commit identifiers stable; replaying produces a linear story that never happened and writes new commits, so it is safe only on a line nobody else has copied.

open as a page

When commit history is rewritten, what happens to the original commits?

level: seniorimportance: should knowfreq 57%

basics

~20 s

Nothing edits them. A rewrite builds new commits carrying the intended content and moves the line's pointer to them; the originals stay in the repository, unreferenced, until a cleanup pass removes them. Anyone holding the old identifiers now has a diverged copy.

open as a page

How do you organise a team's work so that merge conflicts stay rare?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Conflict rate rises with how long two lines stay apart and how much surface each touches, so shrink both: small changes, integrated often. Then remove self-inflicted sources: formatting agreed once and applied mechanically, generated files untracked, modules clearly owned.

open as a page

How can two changes merge with no conflict and still leave the system broken?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Textual merging only guarantees that no two changes touched the same region. Changes in different places can still contradict each other: one side alters what a value means, the other adds a caller assuming the old meaning. Clean merge, broken result.

open as a page

How do you make a shipped release traceable to the exact commit it was built from?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Build every release from a tagged commit in a pipeline, never from a working copy with uncommitted edits, and stamp the version and commit identifier into the built artefact so a running instance can report where it came from.

open as a page

When is hiding unfinished work behind a runtime switch better than hiding it on a branch?

level: seniorimportance: should knowfreq 51%

basics

~20 s

A runtime switch moves the hiding place from version control into the running system: code integrates daily while the behaviour stays off. It trades merge debt for flag debt, meaning more live paths, more states to test, and switches nobody removes.

open as a page

How do you choose a branching model for a team that cannot release its shared line on demand?

level: principalimportance: should knowfreq 38%

basics

~20 s

Derive the model from the obligations rather than adopting a diagram. Establish what actually blocks releasing on demand, such as approval, verification time or a contractual date, and how many versions live in the field, then choose the model paying the least divergence.

open as a page

In version control, what does it mean to sit on a commit that no branch name points to?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

Your position is a specific commit rather than a moving name. Inspecting is fine, but anything you record there gets no name, so once you move away nothing refers to it and the work becomes unreachable and eventually reclaimed.

open as a page

Why can a release tag stored as its own object be signed, while a tag that is only a name pointing into history cannot?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

A signature must cover a fixed sequence of bytes. The object form stores real content (target, name, author, timestamp, message), so there is something to sign; a bare table entry has no content of its own.

open as a page

In a distributed version-control system, what makes one repository the authoritative one?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Nothing in the software. Every copy is structurally a peer, so one copy is authoritative only because the team agrees it is and backs that agreement with access rules on its shared pointers and automation that consumes only that copy.

open as a page