skip to content

How would you create a Git commit using only plumbing commands, without git commit?

level: middleimportance: should knowfreq 30%

answer

  1. four objects, four commands, one ref
  2. content first, names come later
  3. the index is what gets snapshotted
  4. the commit exists before anything points at it
  5. the last step is the one that matters

basics

~20 s

Write the content as a blob with git hash-object -w, put it in the index with git update-index, snapshot the index with git write-tree, wrap that tree with git commit-tree naming a parent, then point a branch at the result with git update-ref.

solid answer

~50 s

Four objects, four commands, in dependency order. `git hash-object -w --stdin` (or on a file) stores the file content as a **blob** and prints its object ID. `git update-index --add --cacheinfo <mode>,<oid>,<path>` records that blob at a path in the index. `git write-tree` turns the whole index into a **tree** object and prints its ID. `git commit-tree <tree> -p <parent> -m "<message>"` creates the **commit** object pointing at that tree, with the given parent, using the identity from `GIT_AUTHOR_*` / `GIT_COMMITTER_*` or your configuration — and prints its ID. At this point the commit exists but nothing references it, so it is unreachable garbage; `git update-ref refs/heads/main <commit>` is what makes it real. Note that none of this touches your working tree — `git read-tree -u` or a checkout would be needed for that — and that `git write-tree` refuses to run while the index has unmerged entries.

code

bash · 5 lines
bash
blob=$(printf 'hello\n' | git hash-object -w --stdin)
git update-index --add --cacheinfo 100644,$blob,greeting.txt
tree=$(git write-tree)
commit=$(git commit-tree $tree -p $(git rev-parse HEAD) -m "add greeting")
git update-ref -m "commit: add greeting" refs/heads/main $commit

go deeper

for a junior

Recall the order — content becomes a blob, the index becomes a tree, the tree becomes a commit, and a ref must then be pointed at that commit.

for a middle

Name each command and what object it produces, and explain why the commit is unreachable until update-ref runs and why the working tree is untouched throughout.

for a senior

Discuss the details that matter in scripting: file modes in update-index, identity and date environment variables, conditional ref updates, and why write-tree fails on unmerged index entries.

for a principal

Use it to frame what porcelain actually contributes — hooks, message policy, signing, safety checks — and therefore what any tooling that bypasses git commit silently loses.

## Why the exercise matters `git commit` looks atomic, but it is a sequence of small object writes plus one ref update. Reproducing it by hand is the clearest possible demonstration that Git is a content-addressed object store with a pointer on top, and interviewers use it exactly for that. The four steps map onto the four things Git stores. ## Step 1 — content becomes a blob ``` blob=$(printf 'hello\n' | git hash-object -w --stdin) ``` `git hash-object` computes the object ID of the given content: hash over the header `blob <size>\0` followed by the bytes. Without `-w` it only computes and prints; `-w` also writes the object into the database. `--stdin` reads from standard input; otherwise you name a file. When hashing a working-tree file, Git applies any configured clean filter and end-of-line conversion unless you pass `--no-filters`, and `--path` lets you tell Git which path's attributes should apply when the content is not coming from that path. The blob knows nothing about a filename. Names live in trees. ## Step 2 — the index gains an entry ``` git update-index --add --cacheinfo 100644,$blob,greeting.txt ``` `--cacheinfo` registers an entry directly from an object ID, without the content needing to exist in the working tree — which is why this works even in an empty directory. The mode is one of Git's small fixed set: `100644` for a regular file, `100755` for an executable, `120000` for a symlink (whose blob contains the target path), and `160000` for a gitlink, the submodule commit pointer. `--add` permits creating an entry that is not already tracked. ## Step 3 — the index becomes a tree ``` tree=$(git write-tree) ``` `git write-tree` serialises the current index into tree objects — one per directory, nested — and prints the ID of the top-level tree. Because trees are content-addressed, writing the same content twice yields the same ID and no new storage. Two properties are worth remembering: write-tree reads the *index*, not the working tree, so unstaged edits are invisible to it; and it fails outright if the index contains unmerged entries, that is, entries at stages other than zero left by a conflicted merge. There is also `git mktree`, which builds a tree from a listing on stdin without involving the index at all. ## Step 4 — the tree becomes a commit ``` commit=$(git commit-tree $tree -p $(git rev-parse HEAD) -m "add greeting") ``` `git commit-tree` creates the commit object: exactly one tree, zero or more parents (`-p` per parent, so two or more make a merge, none makes a root commit), author and committer identity lines, and the message, taken from `-m` or from stdin. Identity comes from `user.name` and `user.email` unless overridden by the `GIT_AUTHOR_NAME`, `GIT_AUTHOR_EMAIL`, `GIT_AUTHOR_DATE`, `GIT_COMMITTER_NAME`, `GIT_COMMITTER_EMAIL` and `GIT_COMMITTER_DATE` environment variables — which is exactly how history-rewriting tools reproduce original timestamps. Note what commit-tree does *not* do: it does not move any ref, does not consult `HEAD` for a parent, does not run hooks, and does not update the working tree. It creates one object and prints its ID. ## Step 5 — a ref makes it reachable ``` git update-ref -m "commit: add greeting" refs/heads/main $commit ``` Until this runs, the new commit is unreachable: nothing under `refs/`, no `HEAD`, no index entry points at it, so it is eligible for pruning once the grace period passes. `git update-ref` writes the new value and, with `-m`, appends a reflog entry. Passing the expected old value as a third argument makes the update conditional, which is the safe form in scripts. `git update-ref --stdin` applies many ref updates as one transaction. ## What git commit adds on top The porcelain wrapper contributes: resolving the parent from `HEAD` automatically; opening an editor and applying `commit.template` and cleanup rules; running `pre-commit`, `prepare-commit-msg`, `commit-msg` and `post-commit` hooks; signing when configured; refusing empty commits; and writing a sensible reflog message. All real value — but none of it is the object model. ## Getting the working tree in sync Because the hand-built commit only touched the object database and a ref, your files on disk are unchanged. `git read-tree -u` (optionally with `--reset`) updates the index and working tree from a tree object; `git checkout` or `git switch` is the porcelain equivalent. Forgetting this is the usual reason the exercise appears not to have worked. ## The one-line summary Content to blob (`hash-object -w`), blob to index (`update-index`), index to tree (`write-tree`), tree to commit (`commit-tree`), commit to branch (`update-ref`). Say that back in order and the object model has been demonstrated.

  • What happens if you run commit-tree but never run update-ref?
    The commit object exists in the database but nothing reaches it — no ref, no HEAD, no index entry, and no reflog entry either, since update-ref is what writes one. It is unreachable from the start and will be pruned once it falls outside the gc grace period. You can still use it by object ID until then.
  • How do you create a merge commit with commit-tree?
    Pass `-p` more than once: `git commit-tree <tree> -p <first> -p <second> -m "merge"`. Parent order is significant — the first parent is the branch being merged into, which is what `HEAD^1` and first-parent history follow. commit-tree does no merging itself; you must supply a tree that already represents the merged result.
  • Why can git write-tree fail during a conflicted merge?
    Because the index then holds unmerged entries, the same path recorded at stages one, two and three for the common ancestor, ours and theirs. A tree object has no way to represent three versions of one path, so write-tree refuses until the conflict is resolved and each path is back at stage zero.
  • How would you set an author date different from the commit time?
    Export `GIT_AUTHOR_DATE` before running `git commit-tree`, and `GIT_COMMITTER_DATE` if the committer timestamp should also be fixed. commit-tree reads both, along with the corresponding name and email variables. That is exactly how history-rewriting tooling preserves original authorship while producing new commit objects.

saying these in an interview costs you the question

  • Thinks commit-tree moves the current branch
  • Believes write-tree reads the working tree
  • Assumes a blob knows its own filename
  • Skips update-ref and wonders where the commit went
  • Expects the working tree to change automatically

context