In Git, what is the difference between loose objects and packfiles, and why pack?
answer
- one file per object at first
- one archive instead of many files
- similar objects stored as differences
- chain depth and window are tunable
- the sidecar file maps IDs to offsets
basics
~20 sGit first writes each new object as its own zlib-compressed loose file under .git/objects. Packing rewrites many objects into a single .pack file that uses delta compression, plus an .idx index for lookup, cutting both disk usage and file count.
solid answer
~50 sWhen you commit, Git writes every new blob, tree and commit as a **loose object**: one file per object under `.git/objects/ab/cdef…`, zlib-compressed on its own, named by its object ID. That is fast to write but wasteful — two near-identical revisions of a big file are stored twice in full, and a large repo would end up with hundreds of thousands of tiny files. `git gc` (or `git repack`) consolidates them into a **packfile**: a single `.pack` under `.git/objects/pack/` holding many objects, where similar objects of the same type are stored as *deltas* against a chosen base rather than in full. Each `.pack` ships with an `.idx` file mapping every contained object ID to its byte offset so lookup stays fast. Packing never changes object IDs or history — it is purely a storage representation, and the same format is used on the wire during fetch and push.
code
console · 9 lines$ git count-objects -v
count: 12
size: 48
in-pack: 4213
packs: 1
size-pack: 9048
prune-packable: 0
garbage: 0
size-garbage: 0go deeper
Recall that Git stores each new object as its own compressed file and later consolidates many of them into a packfile, and that both are inside .git/objects.
Be ready to explain delta compression, why the .idx exists, and that object IDs are unaffected by which representation is used. Naming pack.depth and pack.window shows real familiarity.
Show you can measure it: count-objects -v to see loose versus packed, verify-pack -v to find what dominates the pack, and an understanding that deep delta chains trade disk for CPU on every read.
Own the tradeoff at scale: packing policy affects clone time, CI disk, and server CPU. Be ready to argue when aggressive repacking is worth its one-off cost and when it is cargo cult.
## Two representations of the same object Git's object database is content-addressed: every blob, tree, commit and annotated tag is stored under the hash of its contents. That hash is computed over the object's *logical* content (a header of the form `<type> <size>\0` followed by the bytes), never over how the bytes happen to sit on disk. This is why Git can keep the same object in two completely different physical layouts without anything else in the system noticing. **Loose objects** are the write-path format. When you run `git add` or `git commit`, each new object is deflated with zlib and written as its own file at `.git/objects/<first two hex chars>/<remaining hex chars>`. The two-character directory is just fan-out so no single directory holds a million entries. Loose objects are simple and cheap to create — one file, one write, no coordination with anything already stored — and they are individually compressed but never compressed *against each other*. **Packed objects** are the storage format. A packfile is one big `.pack` file under `.git/objects/pack/` containing many objects concatenated together, accompanied by a `.idx` file. Recent Git versions may also write a `.rev` reverse-index file alongside them. ## What the .idx file does A packfile on its own is not randomly accessible: you cannot know where object `3b18e51…` lives inside a multi-gigabyte file. The `.idx` is a sorted index of every object ID in the pack together with its offset into the `.pack`, plus a 256-entry fan-out table keyed by the first byte of the ID. Looking an object up is a fan-out lookup followed by a binary search — logarithmic, not a scan. The `.idx` contains no object data; delete it and Git can rebuild it from the pack with `git index-pack`. ## Delta compression The real win in a packfile is that objects are not all stored whole. During packing, Git sorts candidate objects heuristically (by type, path name, size) and, for each object, searches a sliding *window* of nearby candidates for a good base. If one is found, the object is stored as a **delta** — a compact instruction stream saying "copy these byte ranges from the base, then insert these literal bytes" — and the result is zlib-compressed on top of that. A base may itself be a delta, forming a **delta chain**. Several rules constrain this. Git only deltifies objects of the same type, so a blob is never delta'd against a tree. Deltas may point at bases either by offset within the same pack or by object ID, but the base must be resolvable — a "thin" pack whose bases live elsewhere is only legal in transit and is completed on receipt. Chain depth is bounded by `pack.depth` (default 50) and the search width by `pack.window` (default 10), because a long chain costs CPU on every read: reconstructing the object means walking back to a whole base and applying each delta in turn. Blobs above `core.bigFileThreshold` (default 512 MiB) are stored undeltified. `git gc --aggressive` mainly widens the search window — `gc.aggressiveWindow` defaults to 250 — while `gc.aggressiveDepth` defaults to 50. A crucial consequence: delta bases have nothing to do with commit order. Git may store an *older* revision of a file as a delta against a newer one, because packing prefers newer objects as whole bases so that recent history reads fast. This is not a per-file version chain the way older centralized systems worked. ## When packing happens Packing is not automatic on commit. It happens when `git gc` runs — explicitly, or via `git gc --auto`, which many commands invoke and which triggers once loose objects exceed `gc.auto` (default 6700) or the number of packs exceeds `gc.autoPackLimit` (default 50). A clone arrives already packed, because the sending side generates a packfile for the transfer. `git repack -a -d` rewrites everything into one pack and deletes the now-redundant ones; `git unpack-objects` does the reverse for a pack fed on stdin. ## Seeing it yourself `git count-objects -v` reports `count` and `size` for loose objects (size in KiB) against `in-pack`, `packs` and `size-pack` for packed ones. `git verify-pack -v` on an `.idx` lists every object with its type, size, packed size and delta chain depth, which is how you find what is actually consuming space. Because the physical layout is invisible above the object layer, none of this affects correctness — it affects clone time, disk footprint, and how fast history traversal reads from disk.
- What limits how long a delta chain can get, and why does the limit exist?`pack.depth` (default 50) caps chain length and `pack.window` (default 10) caps how many candidate bases Git examines. Depth is a read-time cost: reconstructing a deeply chained object means fetching a whole base and applying every delta above it. Deeper chains shrink the pack but slow every read of those objects.
- Can a delta ever be computed between a blob and a tree?No. Git only deltifies objects of the same type, so blobs are delta'd against blobs and trees against trees. Blobs larger than `core.bigFileThreshold` (default 512 MiB) are also stored whole rather than delta'd, which is why one huge binary tends to dominate a repository's size.
- Does packing change any object IDs?No. An object's ID is the hash of its logical type-and-content, not of its on-disk encoding, so the same commit keeps the same ID whether it is loose or packed, whole or stored as a delta. Packing is a pure storage optimisation and is invisible to every command above the object layer.
Loose objects are like keeping every draft of a document as its own separate zip file; a packfile is one archive that stores each draft as the edits needed to reach it from a similar draft, with a table of contents at the front.
saying these in an interview costs you the question
- Thinks packing rewrites history or changes object IDs
- Says a packfile is just the repository zipped once
- Believes deltas form per-file chains in commit order
- Assumes loose objects are stored uncompressed
- Thinks the .idx file holds the object data