skip to content

In Git, why do ten identical copies of a file cost only one stored object?

level: middleimportance: should knowfreq 45%

answer

  1. the key is derived from the value
  2. a blob knows nothing about its filename
  3. a header sneaks in before the bytes
  4. writing an existing object is a no-op

basics

~20 s

Git names objects by a hash of their content, not by path, so identical bytes always produce the same blob id. The ten paths become ten tree entries that all point at one stored blob.

solid answer

~40 s

Git is a content-addressable store: an object's name *is* the hash of its content, so writing the same bytes twice computes the same id and the second write is a no-op. A blob holds only file bytes — no name, no path, no timestamp — so ten identical files anywhere in the tree resolve to one blob, referenced from ten tree entries that differ only in name. The same property makes a pure rename cheap: the blob is untouched and only trees change. Note the hash is not the hash of the raw file: Git prefixes a header of the form `blob <bytesize>` plus a NUL byte before hashing, then zlib-compresses the result into `.git/objects/`, splitting the id so the first two hex characters form the directory name.

code

console · 4 lines
console
$ printf 'hello\n' | git hash-object --stdin
ce013625030ba8dba906f756967f9e9ca394464a
$ printf 'blob 6\0hello\n' | shasum
ce013625030ba8dba906f756967f9e9ca394464a  -

go deeper

for a junior

Recall that Git names objects by hashing their content, so identical files share one stored object and copying or moving a file is cheap.

for a middle

Explain the mechanics: a type-and-length header is hashed ahead of the content, the object is zlib-compressed into a fanned-out path under the object directory, and re-writing an existing id is a no-op.

for a senior

Draw the operational conclusions: content is sticky, so a committed secret or a huge binary stays in history until it is rewritten, and rename tracking is an inference layered on top of dedup.

for a principal

Own the tradeoff at repository scale: content addressing buys integrity and free dedup but makes churn on large binaries expensive, which is what drives policy on what belongs in the repository at all.

## Content addressing In most stores you pick a key and attach a value to it. Git inverts that: the value determines the key. An object's id is a cryptographic hash computed over the object's own bytes, so two identical objects are literally the same object, and there is no way to have two copies of the same content under different ids. Deduplication is not a feature Git implements; it is a consequence of the naming scheme. ## What exactly is hashed Git hashes a header followed by the content. For a file, the header is the word `blob`, a space, the content length in bytes, and a NUL byte. That is why the blob id is *not* the plain hash of the file, and why you can reproduce it by hand: running `git hash-object` on content consisting of `hello` and a newline yields `ce013625030ba8dba906f756967f9e9ca394464a`, and hashing the literal bytes `blob 6`, NUL, `hello` and a newline with a SHA-1 tool yields exactly the same value. Trees, commits and tags are hashed the same way with their own type words. ## Where the object lands A newly written object is zlib-compressed and stored *loose*: one file per object under `.git/objects/`, where the first two hex characters of the id name the subdirectory and the rest name the file. If that path already exists, Git has nothing to do — the content is already there. ## Consequences you should be able to name - **Copies are free.** Ten identical files add one blob; the extra cost is ten tree entries, which are tiny. - **Renames and moves are cheap.** Moving a file writes new trees but reuses the blob untouched. Git does not record renames at all — `git log --follow` and rename detection in diffs *infer* them afterwards by comparing content. - **Unchanged subtrees are reused.** Committing a change deep in one directory reuses every untouched sibling tree object wholesale. - **A blob has no identity of its own.** It does not know its filename, its path, or which commits reference it; that knowledge lives entirely in trees. - **Content is sticky.** Once written, an object stays until nothing reachable references it, which is why accidentally committing a large file or a secret is not fixed by deleting the file in a later commit — the blob is still in history. ## The tradeoff Content addressing gives integrity for free: because ids are derived from bytes, silent corruption of an object changes its content but not the id others use to ask for it, so Git can detect the mismatch. The price is that content, not location, is the unit of storage — a one-byte change to a large file writes a whole new blob, and only later compaction into packfiles reduces that cost with delta compression.

  • Why doesn't a plain SHA-1 checksum of a file match its Git blob id?
    Git hashes a header before the content: the type word, a space, the byte length, and a NUL. Hash that header plus the file bytes and you reproduce the blob id exactly. The header is what keeps the blob, tree, commit and tag id spaces distinct.
  • If content addressing dedupes automatically, how does Git handle a rename?
    It doesn't record one. The blob is reused unchanged and the trees are rewritten, so the rename is invisible in storage. Diff and log rename detection reconstructs it later by matching identical or similar content between the old and new paths.

saying these in an interview costs you the question

  • Thinks the blob id is the plain hash of the file bytes
  • Claims Git stores a filename inside the blob
  • Says Git records renames explicitly in the commit
  • Believes deleting a file in a later commit removes its blob from history

context