In Git, what does a commit object actually store?
answer
- one pointer down, several pointers back
- two people are recorded, not one
- the diff is not in there
- its own bytes decide its name
basics
~10 sA commit stores the hash of one tree (a full snapshot of the tracked files), the hashes of its parent commits, author and committer identity with timestamps, and the message. It stores no diff.
solid answer
~40 sA commit object is a short text record. It names exactly one **tree**, the complete snapshot of tracked files at that point; it lists zero or more **parent** commit hashes; it carries an **author** line (who wrote the change, with a timestamp) and a **committer** line (who created this commit object, with its own timestamp); optionally a signature header; then a blank line and the message. All project state is reached through the tree, so a commit is a snapshot, not a patch — `git log -p` computes diffs on the fly by comparing trees. The commit id is the hash of exactly those bytes, so changing the message, a timestamp, a parent or the tree produces a *different* commit rather than an edited one.
code
console · 6 lines$ git cat-file -p HEAD
tree b4ed918248039b78f24383523fa4e51f80994fac
author Tester <[email protected]> 1787132493 +0400
committer Tester <[email protected]> 1787132493 +0400
add fgo deeper
Recall the field list — tree, parent, author, committer, message — and that a commit is a snapshot of everything tracked, not just the files you touched.
Explain why snapshots are cheap (unchanged files reuse the same blob and tree objects) and why the hash makes a commit immutable rather than editable.
Use the model to reason about operations: rebase, amend and cherry-pick all create new objects, so recovery is about finding the old ids rather than undoing an edit.
Own the consequence for policy: identity and provenance live in the commit's own bytes, so anything the organisation wants to trust — authorship, signatures, traceability — must be part of that record, not metadata bolted on beside it.
## The object it points at Git stores four object types, each named by the hash of its own content: **blob** (the bytes of one file version, carrying no name), **tree** (one directory level, listing mode + name + object hash per entry), **commit**, and annotated **tag**. A commit sits on top: it references one tree, and that tree references blobs and subtrees, so a single commit hash transitively identifies every byte of the tracked project. ## The bytes of a commit `git cat-file -p HEAD` prints the object verbatim — a header block, a blank line, then the message: - `tree <hash>` — exactly one line, always. - `parent <hash>` — none for the very first commit, one for an ordinary commit, two or more for a merge. - `author <name> <email> <epoch> <timezone>` — who wrote the change. - `committer <name> <email> <epoch> <timezone>` — who produced this commit object. - optionally `gpgsig` with an inline signature. That is the whole record. Note what is absent: no diff, no list of changed files, no branch name, and no reference to any child commit. Git derives all of those. ## Snapshot, not delta Every commit points at a *complete* tree. Unchanged files are not duplicated, because a tree entry just reuses the existing blob hash, and an unchanged subdirectory reuses its whole subtree. So a snapshot model costs about the same as a delta model while making checkout a direct lookup instead of a replay of patches. When you ask for a diff, Git compares the commit's tree with its parent's tree and reports the entries whose hashes differ. ## Why the hash is an identity The commit id is the hash over the header plus the message. Because the tree hash is inside those bytes, and the parent hash is too, the id commits to the entire history behind it: change an ancestor and every descendant id changes. This is why history rewriting produces *new* commits rather than modifying existing ones, and why an old commit is not destroyed by an amend — it simply becomes unreferenced. ## Author versus committer Both lines exist because Git was built for patches travelling by mail. The author is the person who wrote the change and is preserved when the commit is replayed elsewhere; the committer is whoever created *this* object, and is updated when a commit is rebased, cherry-picked or amended. Two commits with identical trees and messages therefore differ in id whenever these identities or timestamps differ. ## Interview framing The crisp answer is: "tree, parents, author, committer, message — a snapshot plus provenance." Then add the two consequences interviewers are actually probing for: commits are immutable because the id is a hash of the content, and diffs are computed, not stored.
- If commits are snapshots, why doesn't the repository grow enormously with unchanged files?Because a tree entry for an unchanged file reuses the existing blob hash — the same object, not a copy — and an untouched subdirectory reuses its entire subtree object. Only objects whose content actually changed are new.
- Why do two commits with the same change and message end up with different hashes?The hash covers the whole header, including parent hashes, author and committer identities, and both timestamps. Different parents or different times produce different bytes, hence a different id — which is exactly why a rebased commit is a new object.
- What is the difference between the author and the committer fields?The author is who originally wrote the change and stays fixed as the commit travels; the committer is who created this particular commit object. Rebase, cherry-pick and amend keep the author but rewrite the committer and its timestamp.
saying these in an interview costs you the question
- Says a commit stores a diff or patch of the changes
- Claims the commit records which branch it was made on
- Thinks amend edits the existing commit object in place
- Believes each commit copies every file again
- Cannot distinguish author from committer