skip to content

In Git, what are the working tree, the index, and HEAD?

level: juniorimportance: must knowfreq 82%

answer

  1. count the snapshots Git juggles
  2. one is on disk, two are not
  3. the middle one is a file in .git
  4. HEAD resolves to the current commit

basics

~20 s

Git juggles three trees: the working tree (your files on disk), the index or staging area (the proposed next commit), and HEAD (the commit you are on). git add copies working tree into index; git commit turns the index into a commit.

solid answer

~50 s

Git keeps three snapshots of the same file set. The **working tree** is the checked-out files you edit on disk. The **index** (staging area, cache) is a real binary file, `.git/index`, holding the exact content proposed for the next commit. **HEAD** is a ref that resolves to the commit you are currently on, so "the HEAD tree" means the snapshot that commit records. Commands are movements between those trees. `git add` copies working-tree content into the index. `git commit` writes the index out as tree objects plus a new commit, then advances the branch HEAD points at, so all three match again. `git switch` moves HEAD and rewrites both index and working tree to the target commit. The three flags of `git reset` are named for how far down this stack the move reaches. The diff commands fall out of the model: `git diff` compares working tree to index, `git diff --cached` compares index to HEAD, `git diff HEAD` compares working tree to HEAD.

go deeper

for a junior

Be ready to name the three trees and say which command moves content between which pair. Knowing that git add stages and git commit records the staged snapshot is the screening bar.

for a middle

Explain that the index is .git/index, a full snapshot with blob ids and cached stat data, and derive git status sections and the git diff variants from the model rather than memorising them.

for a senior

Show that you use the model diagnostically: place any unfamiliar command on the diagram, and predict what a teammate's confusing status output means before touching their machine.

for a principal

Frame the index as the reason Git supports deliberate, reviewable commits at all, and connect it to team norms about atomic commits and what a reviewer can actually reason about.

## Why three, not two Most version-control tools have two states: what you have and what is committed. Git inserts a third, the index, between them. That extra buffer is what makes it possible to compose a commit deliberately instead of committing whatever happens to be on disk at that moment. **Working tree.** The ordinary directory you edit, plus a `.git` directory at its root. Nothing here is special to Git; it is just files. Git notices changes only when you run a command that looks. **Index.** A single binary file, `.git/index`. Despite the name "staging area", it is not a list of filenames you queued up: it is a full snapshot. Each entry records a path, a file mode, the object id of the blob holding that path's content, a stage number, and cached `stat` data from the last time Git looked at the file. When you `git add` a file, Git hashes its current content into a blob object in the object database and points the index entry at that blob. The bytes you staged are already durable at that moment, before any commit exists. **HEAD.** A ref, normally a symbolic ref stored in `.git/HEAD` containing something like `ref: refs/heads/main`. Following it gives a branch, which gives a commit, which points at a tree object. "The HEAD tree" is shorthand for that commit's snapshot. In detached-HEAD state, `.git/HEAD` holds a commit id directly. ## What each command moves - `git add <path>` — working tree to index. - `git commit` — index to a new commit; the branch HEAD points at moves forward, so index and HEAD agree again. - `git commit -a` — a shortcut that stages tracked modifications and deletions first; untracked files are still not included. - `git switch <branch>` / `git checkout <branch>` — moves HEAD, then rewrites index and working tree to match, carrying uncommitted changes across only when they do not collide. - `git reset` — moves the branch pointer and, depending on the flag, the index and working tree with it. ## Reading status and diff through the model `git status` is really two comparisons printed together: HEAD versus index ("Changes to be committed") and index versus working tree ("Changes not staged for commit"), plus a third list of paths in neither, the untracked files. The diff family is the same split: bare `git diff` shows index-versus-working-tree, `git diff --cached` (or `--staged`) shows HEAD-versus-index, `git diff HEAD` shows HEAD-versus-working-tree. If a file appears in both status sections at once, you staged one edit and then edited the file again — the index holds the older content and will be what the commit records. ## Consequences worth internalising Because the index is a snapshot rather than a queue, editing a file after staging it does not update what is staged. Because staging writes real blob objects, content you `git add` and then overwrite on disk is still recoverable from the object database. And because a commit is built from the index alone, the committed state may never have existed on disk in exactly that form — which is exactly why you verify a partially staged commit before pushing it.

  • Besides the path and blob id, what does the index cache for each entry, and why?
    Each entry carries stat data from the last time Git looked at the file — mtime, ctime, size, inode, device, uid and gid. `git status` lstats each tracked path and compares; a match means "unchanged" without reading a byte of content. On a mismatch Git reads and hashes the file, and if the hash still equals the recorded blob id it refreshes the cached stat instead of reporting a change.
  • If you stage a file and then edit it again before committing, what gets committed?
    The staged version. The index holds a blob captured at `git add` time, and `git commit` writes the index, so your later edit stays only in the working tree — which is why the file shows up under both status sections. Run `git diff --cached` before committing to see exactly what the commit will contain.
  • During an unresolved merge, how does the index represent a conflicted path?
    With up to three entries for the same path at stage 1, 2 and 3 — the merge base, our side and their side — instead of the usual single stage-0 entry. Staging the resolved file collapses them back into one stage-0 entry, which is what tells Git the conflict is settled.

The index is a shipping box, not a to-do list: git add puts a copy of the item in the box, so editing the item afterwards leaves the boxed copy untouched.

saying these in an interview costs you the question

  • Calls the index just a list of filenames you queued
  • Says git commit commits everything currently on disk
  • Thinks HEAD is a synonym for the current branch name
  • Believes staging is a UI concept with no on-disk form
  • Assumes editing after git add updates what is staged

context