skip to content

How would you keep a large, busy Git repository's object store fast and compact over time?

level: principalimportance: nice to knowfreq 18%

answer

  1. reactive thresholds versus a schedule
  2. incremental beats one giant repack
  3. an index across several packs
  4. retention windows are a chosen policy
  5. measure before tuning anything

basics

~10 s

Replace opportunistic gc --auto with scheduled maintenance: periodic incremental repacking, a multi-pack-index, a written commit-graph, reflog and prune windows chosen deliberately, and measurement via count-objects and verify-pack rather than guesswork.

solid answer

~50 s

The default policy is reactive: `git gc --auto` fires from ordinary commands once loose objects pass `gc.auto` (6700) or packs pass `gc.autoPackLimit` (50), which means a full repack lands unpredictably, in the middle of someone's workday, on the biggest repository. For a large busy repository I move that to a schedule — `git maintenance` with its incremental-repack, commit-graph, loose-objects and pack-refs tasks — so consolidation happens on a known cadence rather than as a surprise. A **multi-pack-index** keeps lookups fast while several packs coexist, so you do not need one giant pack to stay quick, and `git repack --write-bitmap-index` helps where object-set computation dominates. Retention is a policy decision, not a default: `gc.pruneExpire` and `gc.reflogExpireUnreachable` set how long a mistake stays recoverable. I would avoid routine `gc --aggressive`, measure with `git count-objects -v` and `git verify-pack -v`, and treat "stop admitting huge blobs" as the only fix that actually scales.

code

bash · 4 lines
bash
git maintenance start
git maintenance run --task=incremental-repack
git multi-pack-index write
git count-objects -v

go deeper

for a junior

Know that Git cleans up on its own via gc and that a repository can also be maintained on a schedule; the details are not expected at this level.

for a middle

Explain what triggers gc --auto and what a repack actually does, and be able to read git count-objects -v to say whether the problem is loose objects, too many packs, or content.

for a senior

Show you can operate it: move to scheduled maintenance, use a multi-pack-index with incremental repacking, and diagnose with verify-pack rather than reaching for --aggressive.

for a principal

Own the whole tradeoff — cadence and cost of maintenance, retention windows as a recoverability policy, which derived indexes the workload justifies, and the fact that content policy, not packing, is what bounds long-term growth.

## Why the defaults stop being right at scale Git's out-of-the-box maintenance is opportunistic. Many commands call `git gc --auto`, which does nothing until thresholds trip — `gc.auto` (default 6700 loose objects) or `gc.autoPackLimit` (default 50 packs) — and then does a lot at once. With `gc.autoDetach` it runs in the background, but the CPU and I/O still land on the machine that happened to trip the threshold. On a small repository nobody notices. On a repository with millions of objects, a full repack is minutes of heavy work that will happen at an arbitrary moment, possibly on a CI worker with a time budget. The principal-level move is to make maintenance *scheduled and incremental* instead of *reactive and total*. ## Scheduled maintenance `git maintenance run` executes named tasks, and `git maintenance start` registers a schedule with the host's scheduler. The tasks worth understanding: - **incremental-repack** — consolidates smaller packs into fewer, larger ones a batch at a time, maintaining a multi-pack-index, instead of rewriting everything. - **commit-graph** — keeps the derived commit metadata index current so history and reachability commands stay fast. - **loose-objects** — folds accumulated loose objects into packs in bounded batches. - **pack-refs** — collapses thousands of loose ref files into `packed-refs`, which matters once a repository accumulates many branches and tags. - **prefetch** — updates objects from remotes in the background so foreground fetches are small. With these on a cadence, `gc.auto` can be tuned down or effectively neutralised, because the thresholds are no longer the thing driving the work. ## Layout choices **One pack or many?** A single pack maximises delta opportunities and gives one `.idx` to search. But producing it means rewriting everything, which is exactly the cost you are trying to avoid. A **multi-pack-index** (`git multi-pack-index write`) provides a single lookup structure over several packs, so many packs stop being a lookup problem, and incremental repacking becomes viable. `git repack --keep-largest-pack` and `.keep` marker files let you freeze a big historical pack and only churn the recent objects on top of it. **Bitmaps.** `git repack --write-bitmap-index` (or `repack.writeBitmaps`) precomputes reachability bitmaps that make "which objects are reachable from this ref" cheap. That is a serving-side optimisation — it pays off where object-set computation for clones and fetches dominates, and costs write time and disk otherwise. **Aggressive repacking.** `git gc --aggressive` widens the delta search window (`gc.aggressiveWindow`, default 250) at a large one-off CPU cost. It is defensible once after a bulk history change; as a scheduled job it burns cycles for a marginal, non-cumulative gain and should be justified with a before/after measurement, not assumed. ## Retention is a policy, not a default `gc.pruneExpire` (two weeks) and `gc.reflogExpire` / `gc.reflogExpireUnreachable` (90 and 30 days) decide how long a deleted branch or a reset-away commit remains recoverable, and simultaneously how much dead weight the store carries. Shortening them reclaims space and shrinks the window in which any recovery is possible; lengthening them does the reverse. On developer machines I bias toward recoverability; on ephemeral CI checkouts, which are thrown away anyway, the whole question is moot and the right answer is usually to not maintain them at all. ## Measure, do not guess `git count-objects -v` gives loose `count`/`size` versus `in-pack`/`packs`/`size-pack`, which tells you whether you have a loose-object problem, a too-many-packs problem, or a genuine content problem. `git verify-pack -v` on a pack index lists every object with size and delta depth, which is how you discover that four binaries account for most of the repository. `git fsck --unreachable --no-reflogs` shows how much of the store is dead weight waiting on retention windows. Those three commands answer nearly every "why is this repository big" question without speculation. ## The tradeoff to state out loud All of this is compression of a symptom. Delta compression cannot save you from content that does not delta — large binaries, generated artifacts, archives — and no packing strategy makes a repository that admits them behave like one that does not. So the strategy has two halves: keep the store's *representation* efficient with scheduled incremental maintenance and the right derived indexes, and keep the store's *content* small by not admitting objects that will never compress. Presenting only the first half is the classic incomplete answer.

  • Why not just schedule git gc --aggressive weekly?
    Because it widens the delta search window at a large CPU cost for a one-off, non-cumulative gain: the second run finds almost nothing the first did not. It is defensible once after a bulk history change, backed by a before/after `git count-objects -v`. As a recurring job it consumes hours to save little and delays nothing else usefully.
  • When is keeping several packfiles better than consolidating into one?
    When rewriting everything is the dominant cost. A multi-pack-index gives one lookup structure across many packs, so lookup stays fast while repacking proceeds incrementally in batches. Combined with `--keep-largest-pack` or `.keep` files, the historical bulk is frozen and only recent objects churn.
  • How would you decide the prune and reflog retention windows?
    By asking how long a mistake must stay recoverable and what the dead weight costs. Developer machines favour long windows, because the reflog is the undo button for reset and rebase. Ephemeral CI checkouts favour none at all, since the whole clone is discarded. State the recovery window explicitly rather than inheriting defaults silently.

saying these in an interview costs you the question

  • Treats gc --aggressive as routine scheduled maintenance
  • Assumes one giant packfile is always the goal
  • Tunes gc thresholds without measuring first
  • Ignores that huge binaries do not delta at all
  • Shortens reflog retention with no recovery plan

context