skip to content

questions

5

In Git, what does a commit record about its predecessors, and why is history a DAG?

level: juniorimportance: must knowfreq 62%

answer

  1. Look at what a commit object contains
  2. Count the parent lines
  3. Order of parents is not arbitrary
  4. Hash includes the parent IDs
  5. Edges point backwards only

basics

~20 s

Each Git commit stores an ordered list of parent commit IDs: zero for a root commit, one for an ordinary commit, two or more for a merge. Following those pointers backwards yields a directed acyclic graph, not a straight line.

solid answer

~40 s

A Git commit object holds a `tree` (the snapshot), author and committer metadata, a message, and zero or more `parent` lines. The first commit in a repository has no parent, an ordinary commit has exactly one, a merge commit has two, and an octopus merge has more. A commit's ID is a hash over its own content **including** its parent IDs, so a commit can never point forward and can never gain a parent afterwards — changing a parent produces a different commit. That is why history is a **directed acyclic graph**: every edge points from child to parent, and no cycle can form. Branches and tags are just names pointing into that graph, so `git log`, ranges and reachability are all traversals over parent edges.

code

console · 8 lines
console
$ git cat-file -p HEAD
tree 8f4b1c2a9d3e5f6071829304a5b6c7d8e9f01234
parent 3a1b2c3d4e5f60718293a4b5c6d7e8f901234567
parent 9f8e7d6c5b4a39281706f5e4d3c2b1a098765432
author Ada <[email protected]> 1710000000 +0000
committer Ada <[email protected]> 1710000000 +0000

Merge branch 'feature'

go deeper

for a junior

Be ready to say a commit points at a snapshot plus its parent commits, and that a merge commit has two parents. Knowing the arrows point backwards in time is most of the answer.

for a middle

Explain that the commit ID hashes the parent IDs, so the graph cannot contain a cycle and any rewrite produces new IDs for the whole descendant chain.

for a senior

Show how reachability over parent edges underpins release questions, range selection, and garbage collection, and how first-parent ordering keeps merge history readable in production repos.

for a principal

Be able to argue what graph shape a team should aim for — how much branching and merging the DAG should carry before history stops answering the questions people actually ask of it.

## What a commit actually stores A commit in Git is a small object in the object database. Its content is: one `tree` line naming the root tree object (the whole project snapshot at that point), zero or more `parent` lines naming preceding commits, an `author` line, a `committer` line, optionally a signature, then a blank line and the message. `git cat-file -p HEAD` prints it verbatim. Notice what is absent: no diff, and no branch name. A diff is computed on demand by comparing one commit's tree against another's. A branch name lives in a separate ref file and merely points at a commit; the commit does not know which branches contain it. ## Zero, one, or many parents - **Root commit** — no `parent` line at all. Usually there is exactly one root, but a repository can have several (for example after grafting unrelated histories together, or when a branch was started with `git checkout --orphan`). - **Ordinary commit** — exactly one parent: the commit that was `HEAD` when you ran `git commit`. - **Merge commit** — two or more parents. The order matters and is not arbitrary: the **first parent** is the commit you were on when the merge ran, the second is the branch you merged in. A merge of more than two branches at once (an octopus merge) records three or more parents. ## Why "directed" and why "acyclic" Edges are *directed* because a commit records its parents and nothing records its children. Finding the children of a commit is a search: Git must walk the graph from the refs and see who lists it as a parent. That is exactly why `git log` walks backwards in time by default. The graph is *acyclic* by construction, not by a rule someone enforces. The commit ID is a cryptographic hash of the commit's bytes, and those bytes contain the parent IDs. To make commit A a parent of commit B, A's hash must already exist, so A must already be complete. A cycle would require a commit to contain its own hash, which cannot be constructed. An important corollary: because parents are hashed into the ID, rewriting any commit (amend, rebase, filter) necessarily produces a new ID for it **and** for every descendant. History is immutable; "rewriting" means building a new subgraph and moving refs onto it. ## What the graph buys you Everything higher-level is defined in terms of reachability over these edges. "Is this fix in the release?" is "is commit X reachable from tag v2.0?". Range syntax like `main..feature` is a set difference over reachable sets. Garbage collection keeps objects reachable from refs and reflogs and discards the rest. `git log --graph --oneline` draws the DAG in ASCII, which is the fastest way to see the shape of a merge. ## Frequent misconceptions Candidates often say a commit "stores the changes". It stores a full snapshot tree; the storage layer deduplicates unchanged content and packs deltas, but that is a storage detail, not the data model. Another is that a merge commit contains the merged content "twice" — it has one tree like every other commit, and two parents. A third is that a commit knows its branch; it does not, which is why deleting a branch can strand commits that nothing else points at.

  • How many parents can a single commit have, and what produces more than two?
    Zero for a root commit, one for a normal commit, two for a standard merge. Three or more comes from an octopus merge — merging several branches in one `git merge` invocation. Octopus merges are rare because Git refuses them when any of the branches conflict.
  • Why does amending an old commit change the IDs of every commit after it?
    A commit's ID is a hash of its content, which includes its parent IDs. Change anything in a commit and its ID changes; its child now records a different parent, so the child's hash changes too, and so on down the chain. Rewriting always produces a new subgraph.
  • Given a commit, how does Git find its children?
    It does not store them. Git walks backwards from the refs (branches, tags, remote-tracking refs, and the reflog) and collects commits that list the target as a parent. That is why `git log <commit>` shows ancestors instantly but "what came after this" needs a full traversal.

Each commit is a photo of the whole project with a note on the back listing the photo (or photos) it came after — you can always walk backwards, never forwards.

saying these in an interview costs you the question

  • Says commits store diffs rather than snapshot trees
  • Thinks a commit knows which branch it is on
  • Believes history is a linked list, never branching
  • Claims parents of a merge are unordered
  • Thinks amending edits a commit in place

context

open as a page

In Git, what is the difference between HEAD~2 and HEAD^2?

level: middleimportance: must knowfreq 72%

basics

~20 s

In Git, HEAD~2 walks two steps back along the first-parent chain, i.e. the grandparent. HEAD^2 takes one step to the second parent, which exists only on a merge commit; on a non-merge it is an error.

open as a page

In git log, what does main..feature select, and how does main...feature differ?

level: middleimportance: must knowfreq 68%

basics

~20 s

In git log, main..feature lists commits reachable from feature but not from main. main...feature lists the symmetric difference: commits reachable from either side but not both. In git diff, three dots means something else — merge base against the right side.

open as a page

In Git, how do you resolve an expression like HEAD~3 to the exact commit ID?

level: middleimportance: should knowfreq 45%

basics

~20 s

Run git rev-parse HEAD~3. It is Git's revision parser: it turns any gitrevisions expression — branch name, tag, abbreviation, caret/tilde suffix — into the full object ID, printing nothing else, which makes it the safe check before a destructive command.

open as a page

In Git, what does the first parent of a merge commit mean for git log --first-parent?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A merge commit's first parent is the branch that received the merge. git log --first-parent follows only that edge, so it shows one entry per merge and hides the individual commits from merged branches — the integration history of the receiving branch.

open as a page