Why is a version-control commit's identifier derived from its own contents?
answer
- The name comes from the contents
- Change anything, get a different name
- Each name covers its parent's name
- Tamper evidence and one-string comparison
basics
~20 sThe identifier is a hash over the commit's snapshot, parent pointer, authorship metadata and message, so changing anything produces a different identifier. Two copies of history that share an identifier therefore share everything behind it.
solid answer
~40 sWhen a commit is recorded, its whole content - the file snapshot, the parent's identifier, the authorship metadata, the message - is fed through a cryptographic hash function, and the result is the commit's name. Three properties follow. **Nothing is outside the name**: alter one character of the message and the identifier differs. **Identity chains backwards**: because the parent's identifier is an input, a commit's name transitively depends on every commit behind it, so a single identifier vouches for a whole history. **Names need no coordination**: nothing is assigned by a central authority, so two machines that build byte-identical commits agree on the name without communicating. Together these give tamper evidence and a one-comparison equality test between two copies of history - and they are why a commit cannot be edited, only replaced.
go deeper
Know that the commit identifier is computed from the commit's contents rather than handed out in sequence, and that changing anything about the commit - even one character of the message - produces a different identifier.
Explain the mechanics: the hash covers the snapshot, the parent's identifier, the authorship metadata and the message, so identity chains backwards. Be ready to say what that buys - tamper evidence and a one-comparison equality test between two copies.
Show where you rely on the property in production: pinning deploy and incident records to an exact identifier rather than a movable name, and verifying that what is running is what was reviewed. Be equally clear about what it does not prove.
Own the limits. Be ready to discuss authorship verification through signing as a separate control, the migration pressure created by weakening hash functions, and what a provenance policy should require of the teams you are accountable for.
## How the identifier is produced When a commit is recorded, the system feeds the commit's entire content through a cryptographic hash function: the reference to its file snapshot, the identifier of its parent (or parents), the authorship metadata and recorded times, and the message. The fixed-length output becomes the commit's name. Three consequences follow immediately: - **Nothing about the commit sits outside its name.** Change one character of the message, change a recorded time by one second, change one byte in one file, and the identifier is different. - **The name is not assigned.** No server hands it out and no counter increments. Two machines that independently produce byte-identical commits produce the same identifier without ever talking to each other. - **The name covers the past.** Because the parent's identifier is one of the inputs, a commit's name transitively depends on every commit behind it. ## What that buys 1. **Tamper evidence across the whole chain.** Altering a commit from six months ago changes its identifier, which changes its child's, and so on to the newest commit. There is no way to quietly edit the middle of a history and leave the end looking the same. One trusted identifier therefore vouches for everything reachable behind it. 2. **A cheap equality test.** "Do these two copies of the history agree?" becomes a comparison of two short strings rather than a file-by-file comparison of an entire project across every version. This is the property that makes comparing and verifying histories between machines practical at all. 3. **Names that need no coordination.** Any machine can name a commit without asking permission from a central authority, and identical content always produces the identical name. That is what allows many independent copies of a history to talk about the same commits unambiguously. 4. **Immutability by construction.** You cannot edit a commit, because a commit *is* its contents. Producing different contents produces a different commit. Everything that looks like editing history is really creating new commits, and that is a consequence of the naming scheme rather than a rule someone imposed. | | Assigned sequential number | Content-derived identifier | |---|---|---| | Who issues it | A central authority, in order | Nobody - it is computed | | Survives an edit to the record | Yes, silently | No, by design | | Tells you the order of two records | Yes | No | | Works without coordination | No | Yes | | Detects corruption during transfer | No | Yes | The one thing a sequential number is genuinely better at is human ergonomics: "revision 8412" is easier to say than a long hexadecimal string, and it sorts. That is a real loss. The usual mitigation is that a short prefix of the identifier is enough to name a commit unambiguously within one repository. ## What it does not buy - **It is not proof of authorship.** The author fields are ordinary text supplied when the commit is made; anyone can write anyone's name in them. Establishing that a named person really made a commit requires cryptographic signing, which is a separate mechanism layered on top. - **It is not access control or confidentiality.** Anyone who can read the repository can read every historical state it holds. - **It is not a claim that the content is any good** - only that it has not changed since it was recorded. - **It is not permanently strong on its own.** Hash functions weaken as attacks against them improve, and one function long used for this purpose has had practical collision attacks demonstrated against it, which is why the field has been migrating toward stronger functions. The honest framing is: tamper-evident against accident and casual tampering, and only as strong as the underlying function against a determined attacker. ## Where it shows up in real work When a bike-share operations team promotes a build of the dock-availability service, the deployment record stores the identifier of the commit it was built from. Six weeks later - with half the team hired since the deploy and a 2-day median review wait making anyone's recollection unreliable - the question "is what is running the same code we reviewed?" has an exact answer: compare two identifiers. Nobody has to trust a line name, a build label, or somebody's memory. Those are all movable pointers; the identifier is not, which is why incident timelines and deploy records should cite it rather than a name that can be repointed. The same property explains something juniors find surprising. Correcting a typo in the newest commit's message produces a commit with a *different* identifier, and if that commit had already been shared, the shared copy and yours now disagree about what the newest commit on that line is. The identifier changed because the commit changed. Given how the name is produced, it could not have worked any other way. ## What an interviewer is checking The weak answer stops at "it is a hash, so it is unique." The strong answer explains what the uniqueness is *for*: verification without trust, comparison without cost, and immutability that falls out of the naming rather than being enforced by policy - along with a clear-eyed statement of what the identifier does not prove.
- Why can you usually refer to a commit by the first few characters of its identifier?Because a prefix that is unique within one repository identifies the commit unambiguously there. The full identifier is what provides global uniqueness and tamper evidence; the prefix is a convenience for humans, and it stops working the moment two identifiers in that repository share it.
- If the identifier proves the content is unchanged, does it prove who wrote it?No. The author fields are ordinary text supplied when the commit is made, so the identifier only proves that whatever was recorded - including a name someone typed - has not changed since. Establishing that a named person actually made the commit requires cryptographic signing, a separate control layered on top.
- Why should a deployment record cite an identifier rather than the name of a line of development?A line name is a movable pointer: it means a different commit tomorrow. An identifier names one exact state permanently, so weeks later you can answer whether what is running is what was reviewed without trusting anyone's memory or a label that has since moved.
Like a receipt whose number is computed from every line printed on it: you cannot alter an item and keep the number, and each receipt's number also folds in the previous one's.
saying these in an interview costs you the question
- Thinks identifiers are assigned by a central server
- Says a corrected message keeps the same identifier
- Believes the identifier proves who wrote the change
- Assumes the hash guarantees the content is correct
- Cannot explain why the parent's identifier is included