skip to content

questions

27

In Git, what does a commit record about its predecessors, and why is history a DAG?

level: juniorimportance: must knowfreq 62%

answer

  1. Look at what a commit object contains
  2. Count the parent lines
  3. Order of parents is not arbitrary
  4. Hash includes the parent IDs
  5. Edges point backwards only

basics

~20 s

Each Git commit stores an ordered list of parent commit IDs: zero for a root commit, one for an ordinary commit, two or more for a merge. Following those pointers backwards yields a directed acyclic graph, not a straight line.

solid answer

~40 s

A Git commit object holds a `tree` (the snapshot), author and committer metadata, a message, and zero or more `parent` lines. The first commit in a repository has no parent, an ordinary commit has exactly one, a merge commit has two, and an octopus merge has more. A commit's ID is a hash over its own content **including** its parent IDs, so a commit can never point forward and can never gain a parent afterwards — changing a parent produces a different commit. That is why history is a **directed acyclic graph**: every edge points from child to parent, and no cycle can form. Branches and tags are just names pointing into that graph, so `git log`, ranges and reachability are all traversals over parent edges.

code

console · 8 lines
console
$ git cat-file -p HEAD
tree 8f4b1c2a9d3e5f6071829304a5b6c7d8e9f01234
parent 3a1b2c3d4e5f60718293a4b5c6d7e8f901234567
parent 9f8e7d6c5b4a39281706f5e4d3c2b1a098765432
author Ada <[email protected]> 1710000000 +0000
committer Ada <[email protected]> 1710000000 +0000

Merge branch 'feature'

go deeper

for a junior

Be ready to say a commit points at a snapshot plus its parent commits, and that a merge commit has two parents. Knowing the arrows point backwards in time is most of the answer.

for a middle

Explain that the commit ID hashes the parent IDs, so the graph cannot contain a cycle and any rewrite produces new IDs for the whole descendant chain.

for a senior

Show how reachability over parent edges underpins release questions, range selection, and garbage collection, and how first-parent ordering keeps merge history readable in production repos.

for a principal

Be able to argue what graph shape a team should aim for — how much branching and merging the DAG should carry before history stops answering the questions people actually ask of it.

## What a commit actually stores A commit in Git is a small object in the object database. Its content is: one `tree` line naming the root tree object (the whole project snapshot at that point), zero or more `parent` lines naming preceding commits, an `author` line, a `committer` line, optionally a signature, then a blank line and the message. `git cat-file -p HEAD` prints it verbatim. Notice what is absent: no diff, and no branch name. A diff is computed on demand by comparing one commit's tree against another's. A branch name lives in a separate ref file and merely points at a commit; the commit does not know which branches contain it. ## Zero, one, or many parents - **Root commit** — no `parent` line at all. Usually there is exactly one root, but a repository can have several (for example after grafting unrelated histories together, or when a branch was started with `git checkout --orphan`). - **Ordinary commit** — exactly one parent: the commit that was `HEAD` when you ran `git commit`. - **Merge commit** — two or more parents. The order matters and is not arbitrary: the **first parent** is the commit you were on when the merge ran, the second is the branch you merged in. A merge of more than two branches at once (an octopus merge) records three or more parents. ## Why "directed" and why "acyclic" Edges are *directed* because a commit records its parents and nothing records its children. Finding the children of a commit is a search: Git must walk the graph from the refs and see who lists it as a parent. That is exactly why `git log` walks backwards in time by default. The graph is *acyclic* by construction, not by a rule someone enforces. The commit ID is a cryptographic hash of the commit's bytes, and those bytes contain the parent IDs. To make commit A a parent of commit B, A's hash must already exist, so A must already be complete. A cycle would require a commit to contain its own hash, which cannot be constructed. An important corollary: because parents are hashed into the ID, rewriting any commit (amend, rebase, filter) necessarily produces a new ID for it **and** for every descendant. History is immutable; "rewriting" means building a new subgraph and moving refs onto it. ## What the graph buys you Everything higher-level is defined in terms of reachability over these edges. "Is this fix in the release?" is "is commit X reachable from tag v2.0?". Range syntax like `main..feature` is a set difference over reachable sets. Garbage collection keeps objects reachable from refs and reflogs and discards the rest. `git log --graph --oneline` draws the DAG in ASCII, which is the fastest way to see the shape of a merge. ## Frequent misconceptions Candidates often say a commit "stores the changes". It stores a full snapshot tree; the storage layer deduplicates unchanged content and packs deltas, but that is a storage detail, not the data model. Another is that a merge commit contains the merged content "twice" — it has one tree like every other commit, and two parents. A third is that a commit knows its branch; it does not, which is why deleting a branch can strand commits that nothing else points at.

  • How many parents can a single commit have, and what produces more than two?
    Zero for a root commit, one for a normal commit, two for a standard merge. Three or more comes from an octopus merge — merging several branches in one `git merge` invocation. Octopus merges are rare because Git refuses them when any of the branches conflict.
  • Why does amending an old commit change the IDs of every commit after it?
    A commit's ID is a hash of its content, which includes its parent IDs. Change anything in a commit and its ID changes; its child now records a different parent, so the child's hash changes too, and so on down the chain. Rewriting always produces a new subgraph.
  • Given a commit, how does Git find its children?
    It does not store them. Git walks backwards from the refs (branches, tags, remote-tracking refs, and the reflog) and collects commits that list the target as a parent. That is why `git log <commit>` shows ancestors instantly but "what came after this" needs a full traversal.

Each commit is a photo of the whole project with a note on the back listing the photo (or photos) it came after — you can always walk backwards, never forwards.

saying these in an interview costs you the question

  • Says commits store diffs rather than snapshot trees
  • Thinks a commit knows which branch it is on
  • Believes history is a linked list, never branching
  • Claims parents of a merge are unordered
  • Thinks amending edits a commit in place

context

open as a page

In Git, what is HEAD, and how does it differ from a branch name?

level: juniorimportance: must knowfreq 70%

basics

~20 s

HEAD is a pointer to whatever you currently have checked out. Normally it is a symbolic ref holding the name of a branch, such as refs/heads/main, while the branch itself is a ref holding one commit hash.

open as a page

In Git, what is the difference between HEAD~2 and HEAD^2?

level: middleimportance: must knowfreq 72%

basics

~20 s

In Git, HEAD~2 walks two steps back along the first-parent chain, i.e. the grandparent. HEAD^2 takes one step to the second parent, which exists only on a merge commit; on a non-merge it is an error.

open as a page

In git log, what does main..feature select, and how does main...feature differ?

level: middleimportance: must knowfreq 68%

basics

~20 s

In git log, main..feature lists commits reachable from feature but not from main. main...feature lists the symmetric difference: commits reachable from either side but not both. In git diff, three dots means something else — merge base against the right side.

open as a page

In Git, what does a commit object actually store?

level: middleimportance: must knowfreq 82%

basics

~10 s

A commit stores the hash of one tree (a full snapshot of the tracked files), the hashes of its parent commits, author and committer identity with timestamps, and the message. It stores no diff.

open as a page

In Git, what is the difference between loose objects and packfiles, and why pack?

level: middleimportance: must knowfreq 55%

basics

~20 s

Git first writes each new object as its own zlib-compressed loose file under .git/objects. Packing rewrites many objects into a single .pack file that uses delta compression, plus an .idx index for lookup, cutting both disk usage and file count.

open as a page

In Git, which objects does git gc treat as reachable, and what happens to the rest?

level: middleimportance: must knowfreq 50%

basics

~20 s

Git walks out from every ref, HEAD, the index and reflog entries, following commits to their parents, trees and blobs. Anything not reached is unreachable and eventually pruned, but only after a grace period based on object age.

open as a page

What does detached HEAD mean in Git, and how do you keep commits made in that state?

level: middleimportance: must knowfreq 66%

basics

~10 s

Detached HEAD means HEAD holds a commit hash directly instead of naming a branch. Commits still work, but nothing points at them, so create a branch at that commit before switching away.

open as a page

In Git, what makes a repository bare, and why do servers host bare repositories?

level: middleimportance: must knowfreq 55%

basics

~20 s

A bare repository has no working tree and no index: the contents normally found in .git sit directly in the repository directory, with core.bare set to true. Servers use them because nothing should be checked out there, and pushing to a checked-out branch is refused by default.

open as a page

What are the main entries inside a Git repository's .git directory, and what does each hold?

level: middleimportance: must knowfreq 52%

basics

~20 s

HEAD names the current branch, config holds repository-local settings, index is the staging area, objects/ is the object database, refs/ holds branches and tags, logs/ holds the reflogs, hooks/ holds hook scripts, and info/exclude holds unshared ignore patterns.

open as a page

Why can't you commit an empty directory in Git?

level: juniorimportance: should knowfreq 42%

basics

~20 s

Git records files, not directories. A directory exists only as a tree object listing entries, and each entry must name a blob or a subtree, so a folder with nothing in it has nothing to record.

open as a page

In Git, what does git rev-parse HEAD print, and why do scripts rely on it?

level: juniorimportance: should knowfreq 42%

basics

~20 s

It prints the full object ID of the commit HEAD currently resolves to, one line on stdout. Scripts use it because it turns any revision expression into an unambiguous, stable identifier they can record or compare.

open as a page

What happens to your project if you delete the .git directory in a Git repository?

level: juniorimportance: should knowfreq 42%

basics

~20 s

Your files stay exactly as they are on disk, but the folder stops being a Git repository: all history, branches, tags, stashes and configured remotes are gone, because .git is the entire repository and the working tree is only a checkout of it.

open as a page

In Git, how do you resolve an expression like HEAD~3 to the exact commit ID?

level: middleimportance: should knowfreq 45%

basics

~20 s

Run git rev-parse HEAD~3. It is Git's revision parser: it turns any gitrevisions expression — branch name, tag, abbreviation, caret/tilde suffix — into the full object ID, printing nothing else, which makes it the safe check before a destructive command.

open as a page

In Git, why do ten identical copies of a file cost only one stored object?

level: middleimportance: should knowfreq 45%

basics

~20 s

Git names objects by a hash of their content, not by path, so identical bytes always produce the same blob id. The ten paths become ten tree entries that all point at one stored blob.

open as a page

In Git, how does an annotated tag differ from a lightweight tag?

level: middleimportance: should knowfreq 58%

basics

~20 s

A lightweight tag is only a ref under refs/tags pointing straight at a commit. An annotated tag creates a real tag object holding a tagger, a date, a message and an optional signature, and the ref points at that object.

open as a page

How would you create a Git commit using only plumbing commands, without git commit?

level: middleimportance: should knowfreq 30%

basics

~20 s

Write the content as a blob with git hash-object -w, put it in the index with git update-index, snapshot the index with git write-tree, wrap that tree with git commit-tree naming a parent, then point a branch at the result with git update-ref.

open as a page

In Git, how do you inspect an object's type and contents with git cat-file?

level: middleimportance: should knowfreq 35%

basics

~20 s

git cat-file -t prints an object's type, -s its size, and -p pretty-prints its contents according to that type. git cat-file -e tests existence through the exit code, and --batch modes stream many objects for scripts.

open as a page

In Git, what is the difference between porcelain and plumbing commands?

level: middleimportance: should knowfreq 40%

basics

~20 s

Porcelain commands are the human-facing UI — add, commit, log, merge — whose output is formatted for people and may change. Plumbing commands are low-level building blocks like cat-file, hash-object and rev-parse, with stable output meant for scripts.

open as a page

In Git, what does the first parent of a merge commit mean for git log --first-parent?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A merge commit's first parent is the branch that received the merge. git log --first-parent follows only that edge, so it shows one entry per merge and hides the individual commits from merged branches — the integration history of the receiving branch.

open as a page

After removing a huge blob from Git history, why does .git stay large, and what shrinks it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The blob is still reachable from something you forgot: leftover backup refs, tags, remote-tracking refs, other worktrees, or the reflog. Until every path to it is gone and gc prunes past its grace period, the object stays on disk.

open as a page

In Git, what does git ls-files --stage show that git ls-tree HEAD does not?

level: seniorimportance: should knowfreq 22%

basics

~20 s

git ls-tree HEAD reads a committed tree object, while git ls-files --stage reads the index — so it shows staged additions and deletions HEAD lacks, plus a stage number per entry that exposes the three conflicting versions during a merge.

open as a page

In Git, when is the repository directory not the .git folder beside your files?

level: seniorimportance: should knowfreq 32%

basics

~20 s

When .git is a file containing a gitdir: line — used by linked worktrees, submodules and git init --separate-git-dir — or when GIT_DIR and GIT_WORK_TREE, or the --git-dir and --work-tree options, point Git at a repository stored elsewhere.

open as a page

What changes when a Git repository is created with --object-format=sha256?

level: seniorimportance: nice to knowfreq 18%

basics

~10 s

Every object id becomes a 64-hex-character SHA-256 hash instead of a 40-character SHA-1, recorded in the repository config as an extension. The choice is repository-wide, fixed at creation, and cannot interoperate with SHA-1 repositories.

open as a page

What is Git's commit-graph file, and which operations does it make faster?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

The commit-graph is a derived binary cache under .git/objects/info that stores each commit's parents, root tree and date plus generation numbers, so Git can answer history and reachability queries without decompressing commit objects from packs.

open as a page

In Git, why might a branch exist with no file for it under .git/refs/heads?

level: seniorimportance: nice to knowfreq 22%

basics

~10 s

Because refs can be stored two ways. Git compacts loose one-file-per-ref storage into a single .git/packed-refs file, so a branch may exist only as a line there. Loose files, when present, take precedence.

open as a page

How would you keep a large, busy Git repository's object store fast and compact over time?

level: principalimportance: nice to knowfreq 18%

basics

~10 s

Replace opportunistic gc --auto with scheduled maintenance: periodic incremental repacking, a multi-pack-index, a written commit-graph, reflog and prune windows chosen deliberately, and measurement via count-objects and verify-pack rather than guesswork.

open as a page