What changes when a Git repository is created with --object-format=sha256?
answer
- it is declared in config, not inferred
- 40 characters become 64
- every object embeds its neighbours' hashes
- old clients must fail loudly, not guess
basics
~10 sEvery object id becomes a 64-hex-character SHA-256 hash instead of a 40-character SHA-1, recorded in the repository config as an extension. The choice is repository-wide, fixed at creation, and cannot interoperate with SHA-1 repositories.
solid answer
~40 s`git init --object-format=sha256` selects the SHA-256 object format. Git then sets `core.repositoryformatversion` to 1 and `extensions.objectFormat` to `sha256`, so older clients that do not understand the extension refuse the repository rather than misreading it. Object ids become 64 hex characters, and every place a hash appears — tree entries, commit `tree` and `parent` lines, tag objects, refs, the index — uses the new width. The format is a property of the whole repository chosen at creation, and there is no conversion or interoperability with SHA-1 repositories, because each object embeds the hashes of the objects it references, so changing the hash rewrites every id in the graph. Meanwhile Git's default SHA-1 implementation detects the known collision-attack pattern and refuses to proceed.
go deeper
Know only that Git object ids are content hashes, historically SHA-1 and 40 hex characters, and that a newer SHA-256 format exists.
Explain what widens — every id in trees, commits, tags, refs and the index — and that the format is declared in the repository config so incompatible clients refuse it.
Reason about migration: the Merkle structure makes every id change, external references to old ids break, and there is no interoperability, so this is a repository-lifecycle decision rather than a setting.
Own the strategy: weigh the cost of invalidating every recorded commit id across tickets, audits and artefacts against the risk being mitigated, and be honest that ecosystem readiness, not Git itself, gates the move.
## What the flag does `git init --object-format=sha256` creates a repository whose object database is named with SHA-256 instead of SHA-1. Two things are recorded in the local config: `core.repositoryformatversion = 1` and `extensions.objectFormat = sha256`. The version bump matters — a Git that does not recognise an extension it finds in a version-1 repository refuses to operate on it, which is far better than misreading it. The choice is made at creation and applies to the whole repository; there is no per-branch or per-object mixing. ## What visibly changes Object ids grow from 40 hex characters to 64. Everything that carries an id widens with them: tree entries, the `tree` and `parent` lines of commits, the `object` line of tag objects, ref files, the reflog, the index. Hashing still follows the same recipe — a header of type and byte length, a NUL, then the content — so the same file yields a valid but entirely different id under the new algorithm. ## Why conversion is hard This is the part interviewers are testing. Git's data structure is a Merkle graph: a tree names its blobs by hash, a commit names its tree and parents by hash, a tag names its target by hash. Rehashing one object changes its id, which changes every object that referenced it, all the way to the branch tips. So converting a repository is not a re-index; it is a full rewrite in which every id in existence changes, and every id that lives *outside* the object graph — in commit messages, issue trackers, build logs, deployment records, signed artefacts, other people's clones — no longer resolves. Interoperability between the two formats is not implemented, so a SHA-256 repository and a SHA-1 repository cannot exchange history directly. ## Why the sky has not fallen for SHA-1 Git's use of SHA-1 was never a signature scheme by itself, and a practical attack would require producing a colliding object that is also plausible content an unsuspecting maintainer accepts. Recent Git versions additionally hash with a collision-detecting SHA-1 implementation that recognises the byte patterns produced by the known collision technique and aborts rather than writing the object. That is a mitigation, not a fix, which is exactly why the SHA-256 format exists as the long-term path. ## What to say in an interview Make three points: the format is repository-wide and declared through a config extension so old clients fail loudly; the hash is embedded transitively in every object, so migration means rewriting the entire graph and invalidating every id anyone has ever written down; and the SHA-1 story is mitigated in the meantime by collision detection. Do not claim a specific hosting provider or tool supports the format — the honest answer is that ecosystem support, not Git's own implementation, is the limiting factor.
- Why can't Git simply convert an existing repository in place to SHA-256?Because ids are transitive: a tree names blobs by hash, a commit names its tree and parents by hash. Rehashing anything changes every descendant id, so conversion is a complete rewrite of history — and every id recorded outside the repository, in messages, tickets and logs, stops resolving.
- How does an older Git client behave when it meets a SHA-256 repository?The repository declares `core.repositoryformatversion = 1` plus an `extensions.objectFormat` entry. A client that does not understand that extension refuses to operate on the repository instead of misinterpreting object ids — failing loudly is the deliberate design.
saying these in an interview costs you the question
- Claims you can flip an existing repository to SHA-256 with a config setting
- Says SHA-1 and SHA-256 repositories can push and fetch between each other
- Thinks the object id is a hash of the raw file only, unaffected by references
- Asserts Git ignores the repository format version