Why is a version-control commit a full snapshot rather than a stored diff?
answer
- A state, not a change
- It also points backwards
- Differences are computed, not stored
- Whole file set plus parent pointer
basics
~20 sA commit records the complete state of the file set at one moment, plus a pointer to the commit it was built on and who made it and why. Differences between two commits are computed on demand, never stored.
solid answer
~40 sA commit holds four things: a **complete snapshot** of the file set at that moment, a **pointer to its parent** (the commit it was built on top of), **authorship metadata**, and a **message** saying why. The instinct most people arrive with is that a commit stores the lines that changed; the dominant modern model stores states and computes differences on request. Snapshots are not expensive because the snapshot is a listing that references stored content - a file whose contents did not change is stored once and referenced by every snapshot that includes it. The advantage is that reconstructing any historical version is a direct read rather than a replay from a base, and a diff becomes a *view* over two snapshots rather than the storage format itself.
code
pseudocode · 7 linescommit 7c41e9d2
snapshot -> every file in the project at this moment
parent -> a1f0b83c
author -> who made the change, and when
message -> why the change was made
# the identifier 7c41e9d2 is computed from all four entries abovego deeper
Recall that a commit records the whole file set at a moment plus a pointer to what came before, and that it also carries who made it, when, and why. Do not describe it as a list of changed lines.
Explain snapshot versus stored difference: states are stored and differences are computed on request. Be ready to say why whole snapshots are not expensive, and what the parent pointer establishes that a timestamp cannot.
Demonstrate that you reason about history as a graph with reachability rather than as a log of edits. An interviewer expects you to move fluently between what is stored and what is displayed, and to name which everyday operations that distinction makes cheap.
Be prepared to discuss where the model's costs actually land at scale - large binary content, very long histories, a repository shared by many teams - and how you would decide whether a storage or workflow change is worth the disruption it causes.
## What one commit records A commit is a record of the project at a single moment. Four things go into it: - **A complete snapshot of the file set** - every file that exists at that moment with its contents, not only the ones that changed. - **A pointer to its parent** - the commit this one was built on top of. A commit that brings two lines of development together names both; the very first commit in a repository names none. - **Authorship metadata** - who made the change, and when, as recorded by the machine that made it. - **A message** - the human explanation of why the change exists at all. The record is then named by an identifier derived from all of the above, which is what makes it immutable. ## Snapshot, not stored difference The instinct most people arrive with is that version control stores differences: each commit holds "the lines that changed", and the current state is whatever you get by applying all of them in order. Some systems have genuinely worked that way. The dominant modern model does the opposite - it stores states, and computes differences when someone asks for one. | | Snapshot model | Stored-change model | |---|---|---| | What a commit holds | The whole file set at that moment | The change relative to a base | | Reconstructing an old version | Read that commit's snapshot directly | Replay recorded changes from a base forward | | Producing a difference between two commits | Computed on demand from two snapshots | Read out of the stored records | | Cost as history grows | Flat - any version is one lookup | Grows with the length of the chain | | Showing the same change against a different base | Recompute, no stored data to fix up | The stored change is tied to its base | The payoff is that the expensive operation - reconstructing an arbitrary historical version - becomes a direct read, while the cheap operation - showing a difference - happens only when someone asks. It is also why the *presentation* of a change can be tuned freely: ignoring whitespace, detecting a moved file, adjusting how similar two files must be to count as a rename. All of that is a rendering decision applied at read time, because nothing about the difference was frozen into storage. ## Why whole snapshots are not expensive "A snapshot per commit" sounds like copying the entire project thousands of times. It is not, because a snapshot is a **listing** - names mapped to content - and the content it names is stored once and referenced by every listing that includes it. A file untouched for two years is one stored item referenced by every snapshot since. A commit that edits one file in a large project adds roughly that file's worth of new content plus a small amount of structure. Carrying this distinction - a snapshot references content, it does not contain a copy of it - is what separates a candidate who has read the model from one who is guessing. ## What the parent pointer buys The parent pointer is the half of the answer candidates most often leave out, and it does at least as much work as the snapshot. 1. **Ordering that does not depend on clocks.** Machines disagree about the time - clocks drift, get set wrong, or are changed deliberately. The parent link states which commit came first as a fact of construction, not as an observation of a timestamp. 2. **A graph, not a list.** Because a commit can name more than one parent, history is a directed graph rather than a sequence. Any answer that cannot accommodate that is incomplete. 3. **Reachability.** "Everything behind this commit" is a well-defined set: follow parents until you run out. That set is exactly what a single identifier stands for. 4. **A chain of identity.** Each commit's identifier is computed over its parent's identifier, so the name of the newest commit transitively depends on every commit before it. ## Reading a commit in practice When the newly-hired half of a bike-share operations team asks why the dock-availability service rounds its rebalancing window to 37 minutes, the useful artefact is not the current file - it is the commit that introduced the constant. Its snapshot shows the state the author left the code in, its parent shows the state they started from, and its message says why 37 rather than 30. Three commits later the number may have moved twice; because each of those states was stored rather than derived, each one is directly readable instead of something you reconstruct by replaying a chain. ## Misreadings to avoid - **"A commit stores my changes."** It stores a state. Your changes are the difference between it and its parent, computed when someone asks. - **"History is a list."** It is a graph, and the parent pointer is what makes it one. - **"The identifier is a version number."** It is derived from the commit's contents, so it carries no ordering at all - two identifiers tell you nothing about which came first. - **"A commit records the files I had open."** It records what was selected into it, which is a decision, not an observation of your editor.
- If commits are snapshots, how does the system show you what a commit changed?It compares that commit's snapshot with its parent's and reports the differences. Because the comparison happens at read time, the same commit can be shown against different bases, and the rendering can ignore whitespace or detect a moved file without altering anything that was stored.
- Why does a commit not rely on its timestamp for ordering?Timestamps come from whichever machine recorded the commit, and machines disagree - clocks drift, are set wrong, or are changed deliberately. The parent pointer states ordering as a structural fact: this commit was built on that one. Sorting history purely by time can therefore produce an order that contradicts how it was actually built.
- Does a full snapshot per commit make storage grow with project size times commit count?No. A snapshot is a listing that references stored content, and content that did not change is stored once and referenced by every snapshot including it. A commit touching one file adds about that file's worth of new content plus a little structure, not a copy of the project.
A commit is a photograph of the whole room with a note saying which photograph came before it - not a list of the furniture you moved.
saying these in an interview costs you the question
- Says a commit stores only the changed lines
- Thinks a snapshot means copying the whole project
- Cannot name anything a commit holds besides files
- Describes history as a straight list only
- Treats the commit identifier as a version number