skip to content

In Git, how do you inspect an object's type and contents with git cat-file?

level: middleimportance: should knowfreq 35%

answer

  1. read side of the object database
  2. flags for type, size, contents
  3. one flag answers only through exit status
  4. batch modes stream many lookups
  5. shows stored bytes, not checkout output

basics

~20 s

git cat-file -t prints an object's type, -s its size, and -p pretty-prints its contents according to that type. git cat-file -e tests existence through the exit code, and --batch modes stream many objects for scripts.

solid answer

~50 s

`git cat-file` is the read side of Git's object database. Given an object ID or any expression that resolves to one, `-t` prints the type (`blob`, `tree`, `commit` or `tag`), `-s` prints the size in bytes, and `-p` pretty-prints the content in a form appropriate to that type: a commit shows its `tree`, `parent`, `author` and `committer` headers followed by the message; a tree lists mode, type, object ID and name per entry; a blob is dumped as raw bytes. `-e` prints nothing and just exits zero if the object exists, which is the plumbing way to probe. For scripts there are `--batch` and `--batch-check`, which read object names on stdin and stream results, and `--batch-all-objects`, which iterates the whole store. Importantly, `-p` on a blob gives you the bytes as stored — no smudge filter and no end-of-line conversion — so it shows you what Git really has, not what checkout would write.

code

console · 9 lines
console
$ git cat-file -t HEAD
commit
$ git cat-file -p HEAD
tree 4b825dc642cb6eb9a060e54bf8d69288fbee4904
parent 6f1c2ab3d4e5f60718293a4b5c6d7e8f90a1b2c3
author Ada <[email protected]> 1717171717 +0000
committer Ada <[email protected]> 1717171717 +0000

add parser

go deeper

for a junior

Recall that this command lets you look at any Git object: its type with -t and its contents with -p, including printing a commit as plain text.

for a middle

Explain the per-type pretty-printed output and be able to walk commit to tree to blob by hand, naming what each object contains at every step.

for a senior

Show diagnostic use: reading stored bytes to settle filter or line-ending disputes, using -e for existence checks, and streaming with --batch-check instead of process-per-object.

for a principal

Treat it as the ground truth tool: when tooling, filters or hosting behaviour disagree, reading the object itself is the arbiter and should be the first step in any dispute.

## The command's job Everything Git stores is an object identified by the hash of its content. `git cat-file` is how you look at one from the outside, and it is the single most useful command for proving to yourself — or to an interviewer — that you understand the store rather than the UI. It takes an object name. That can be a raw object ID, an abbreviation of one, or any expression that resolves to an object, which is why `git cat-file -p HEAD` works directly. ## The four basic modes - **`-t` (type)** prints one of `blob`, `tree`, `commit`, `tag`. This is the fastest way to discover what you are holding when all you have is a hash. - **`-s` (size)** prints the object's size in bytes — the logical content size, not the compressed on-disk size, so it is unaffected by whether the object is loose or packed or stored as a delta. - **`-p` (pretty-print)** prints the content, formatted per type. For a **commit** you see the exact stored text: a `tree` line, zero or more `parent` lines, `author` and `committer` lines with name, email, epoch timestamp and timezone offset, optional headers such as `gpgsig`, a blank line, then the message. For a **tree** you get one line per entry: mode, type, object ID, tab, name. For an **annotated tag** you see `object`, `type`, `tag`, `tagger` and the message. For a **blob** you get the raw bytes. - **`-e` (exists)** prints nothing and exits zero if the object exists and is valid, non-zero otherwise. It is a test, not a printer. You can also pass an explicit type as the first argument (`git cat-file blob <oid>`) to assert what you expect. ## Reading a commit is the interview moment `git cat-file -p HEAD` is the demonstration that a commit is just a small text object naming one tree and its parents. From there `git cat-file -p HEAD^{tree}` lists the top-level tree, and following an entry's object ID with another `-p` walks down into a subtree or prints a file's contents. Doing that walk out loud — commit to tree to subtree to blob — answers half the object-model questions an interviewer has queued up. ## Raw bytes, not checkout output A subtlety worth stating: `-p` on a blob prints the object's stored content. If the repository uses a clean/smudge filter or end-of-line conversion, the file in your working tree may differ from the blob. cat-file offers `--filters` to apply the checkout-side filters and `--textconv` to apply a configured textconv driver, but the default is deliberately raw. That makes cat-file the right tool for answering "what is actually stored", which is exactly the question in a line-ending or filter dispute. ## Batch modes for scripts Invoking Git once per object is expensive. `--batch-check` reads object names on stdin and, for each, prints `<oid> <type> <size>`; `--batch` prints that header line followed by the object's content. Because the process stays alive, a script can stream thousands of lookups through one invocation. `--batch-all-objects` walks every object in the database instead of reading names from stdin, and combined with `--batch-check` it is the standard way to enumerate and size everything the repository holds. There is also a `--batch-command` mode in recent Git for interleaving different requests on one stream. ## Where it sits among the other plumbing cat-file reads objects; `git hash-object -w` writes them. `git ls-tree` is a specialised, script-friendlier reader for tree objects, with `-r` to recurse and `--name-only` to strip everything but paths — use it when you want tree entries, and cat-file when you want to see the object as it is. `git rev-parse` resolves the expression you feed to either. `git verify-pack` tells you about physical storage. cat-file deliberately says nothing about how the object is stored — loose or packed, whole or delta'd — because at the object layer that distinction does not exist. ## A useful habit When anything about Git surprises you, resolve the object with `git rev-parse` and look at it with `git cat-file -p`. Almost every confusing behaviour — a mode change you did not expect, a submodule entry, a symlink stored as a blob containing a path, a commit whose parent order explains a diff — becomes obvious the moment you read the object itself.

  • How does git cat-file -p on a tree differ from git ls-tree?
    They show the same entries — mode, type, object ID and name. `git ls-tree` is the specialised reader with script-facing options: `-r` to recurse into subtrees, `-t` to include the tree entries themselves, `--name-only`, and `-z` for NUL-terminated output. Use ls-tree when you want to process entries, cat-file when you want to read the object as stored.
  • Why might git cat-file -p on a blob differ from the file in your working tree?
    Because it prints the stored bytes with no checkout-side processing. If a clean/smudge filter or end-of-line conversion is configured, the working-tree file is the converted form while the blob is what Git committed. `--filters` applies the checkout-side filters if you want to compare the two directly.
  • When would you reach for --batch-check instead of running cat-file per object?
    Whenever you are inspecting many objects. `--batch-check` keeps one process alive, reads object names on stdin, and prints object ID, type and size per line, avoiding thousands of process startups. Paired with `--batch-all-objects` it enumerates the whole store, which is how you find what is actually taking up room.

saying these in an interview costs you the question

  • Thinks cat-file shows diffs rather than stored objects
  • Believes -p on a blob applies checkout filters
  • Says commit objects contain the file contents
  • Runs one process per object in a large loop
  • Confuses object size in bytes with on-disk compressed size

context