Why does git rebase give every replayed commit a new SHA even when the content is identical?
answer
- Identity covers more than the file snapshot
- Ancestry is part of what is hashed
- The change cascades down the chain
- One timestamp is rewritten, one is preserved
- The originals are unreferenced, not deleted
basics
~20 sA commit's hash is computed over its whole content, which includes its parent, its message, and both author and committer identity and timestamp. Rebase re-creates each commit with a different parent and a fresh committer timestamp, so the hash must differ.
solid answer
~50 sGit names a commit by hashing the commit object itself, and that object contains far more than the file snapshot: the tree, the parent hashes, the author name/email/date, the committer name/email/date, and the message. Rebase replays commits onto a different base, so the very first replayed commit already has a **new parent**, which changes its hash; that new hash becomes the parent of the next one, so the change cascades through the entire replayed range. On top of that, replay sets a fresh **committer** timestamp while preserving the original author fields — which is why `git log` still shows the original dates but the commits are demonstrably new objects. Even a rebase that changes no file content at all produces new commits. The old ones are not deleted; they stay in the object database, reachable through the reflog, until garbage collection prunes them.
go deeper
Know the fact and one consequence: rebased commits are new commits with new hashes, which is why a pushed branch stops fast-forwarding.
Explain what goes into a commit's hash, why a changed parent cascades through the whole replayed range, and which metadata is preserved versus rewritten.
Connect it to operations: what breaks for teammates and for anything that recorded the old hashes, and how the reflog makes an unwanted rebase fully recoverable.
Own the systemic cost of identity churn — external references to commit hashes, signature invalidation, and where the repo should draw the line on rewriting.
## What Git actually hashes A commit's identifier is a hash of the commit object's bytes. Those bytes are, roughly: a pointer to the **tree** (the snapshot of the whole directory structure), zero or more **parent** pointers, an **author** line (name, email, timestamp), a **committer** line (name, email, timestamp), and the commit **message**. Change any byte in any of those fields and you get a different identifier. This is what makes Git history tamper-evident — a commit's name commits to everything about it, including its ancestry. The important implication for rebase: identical file content is not enough for an identical hash. Two commits with the same tree but different parents are different objects with different names. ## Why rebase must produce new objects Rebase's whole purpose is to change the base a series of commits sits on. The first commit in the replayed range is re-created with `main`'s current tip as its parent instead of the old fork point. That alone changes its bytes, and therefore its hash. The second commit is then re-created with the *new* first commit as its parent — so it changes too, regardless of whether anything else about it changed. The effect cascades: **every commit from the rebase point forward gets a new identity**, and so does every commit that descends from them. This is not an implementation quirk that a cleverer rebase could avoid. Making commit identity depend on ancestry is the design; a commit that kept its name while changing its parent would break the integrity property the whole model rests on. ## Author versus committer Rebase preserves the **author** fields — the person who originally wrote the change and when — and rewrites the **committer** fields to whoever is running the rebase, at the current time. This is why a rebased branch still shows original dates in `git log`'s default output while `git log --format=%cd` (committer date) shows today. It also means that even a hypothetical rebase where the parent happened to be unchanged would still yield a new hash, because the committer timestamp moved. ## What happens to the originals Rebase does not modify or delete the old commits — commits are immutable, so it cannot. It writes new objects and then moves the branch ref to the last new commit. The old chain is still in the object database; it is simply no longer reachable from the branch. It remains reachable from the reflog, which records the branch's previous positions, so `git reflog` plus a `git reset --hard` to the recorded hash restores the pre-rebase state exactly. Only once reflog entries expire and garbage collection runs do the unreferenced objects actually go away. ## The consequences people care about - **Force-push.** The remote still points at the old chain. Your new tip is not a descendant of it, so the push is not a fast-forward and is rejected unless forced. - **Teammates diverge.** Anyone holding the old commits now has commits that no longer exist upstream, and a naive pull produces a merge of the old and new copies — the duplicated-commits mess that follows an unannounced rebase. - **External references rot.** Anything that recorded the old hashes — build records, links, notes, a `git bisect` session in progress — refers to commits that are no longer on the branch. - **Signatures.** A signature covers the commit object; re-creating the object invalidates any signature it carried unless the rebase re-signs. ## Detecting equivalence anyway Even though the hashes differ, Git can tell that a replayed commit introduces the same change as its original: it computes a **patch-id**, a hash of the change itself rather than the commit object, and uses it to notice duplicates. That is how a rebase can skip commits already present upstream and report that a commit became empty. So "the SHA is different" and "Git cannot tell they are the same change" are two separate statements — the first is true, the second is not. ## The takeaway to state in an interview Rebase copies rather than moves, because identity includes ancestry. Almost every practical rebase question — why force-push, why teammates suffer, why the originals are recoverable — is a corollary of that one sentence, and being able to derive them on the spot is what distinguishes a mechanical understanding from a memorised list.
- Which commit metadata does rebase preserve, and which does it rewrite?It preserves the author name, email and date — the original writer and when they wrote it. It sets the committer name, email and timestamp to whoever ran the rebase, now. That is why default log output still shows old dates while the commits are demonstrably new objects.
- If the hashes all change, how can Git tell a replayed commit is the same change as the original?By patch-id: a hash computed over the change itself rather than over the commit object. Git uses it to notice that a commit's change is already present upstream, which is how a rebase can drop commits that became empty or that were already applied.
- Where do the pre-rebase commits go, and how would you get back to them?Nowhere — commits are immutable, so rebase writes new objects and moves the branch ref. The old chain stays in the object database, unreferenced by the branch but recorded in the reflog. Reading the branch's previous position from the reflog and resetting to that hash restores the pre-rebase state exactly.
saying these in an interview costs you the question
- Thinks the hash covers only the file contents
- Says the old commits are deleted immediately
- Claims a rebase that changes no files keeps the hashes
- Believes only the first replayed commit changes identity
- Cannot connect new hashes to the force-push requirement