In Git, how do you tell whether a commit was already cherry-picked upstream?
answer
- identity comparison cannot work here
- compare the change, not the commit
- a hash of the diff, offsets ignored
- one command prints plus and minus signs
- the same trick makes rebase drop duplicates
basics
~20 sCompare by patch content, not by SHA. git cherry <upstream> <head> marks each commit - when an equivalent change already exists upstream and + when it does not, using a patch-id — a hash of the diff that ignores line numbers and whitespace context.
solid answer
~40 sA cherry-picked commit has a different SHA from its original, so identity comparison is useless. Git instead computes a **patch-id**: a hash of the commit's diff, normalised so that context line numbers and whitespace do not affect it. Two commits with the same effective change share a patch-id. `git cherry <upstream> [<head>]` uses this to list your commits with `+` for "not upstream" and `-` for "an equivalent patch is already there", which is the direct answer to "has this been backported yet". `git log` exposes the same machinery through `--cherry-mark` and `--cherry-pick` when comparing two branches with a symmetric difference. The same detection is why rebase quietly drops commits already applied upstream. It is a heuristic: a change that was adjusted while being picked has a different patch-id and will not match.
code
console · 4 lines$ git cherry -v release-1.4 main
+ 9f2a1c3 add retry to session refresh
- 4b7e0aa fix null deref in session lookup
+ c11d902 tighten rate limiter defaultsgo deeper
Know that a cherry-picked commit has a different SHA, so Git needs a way to compare the change itself rather than the commit id.
Explain the patch-id — a normalised hash of the diff — and name a command that reports which commits already have an equivalent on the other branch.
Use it operationally to audit what still needs backporting between release lines, and state its limits when a pick was adjusted or split.
Decide how backport state is tracked across many release branches — content detection, recorded provenance, or an external record — and what accuracy the process actually requires.
## Why SHAs cannot answer the question Cherry-picking copies a change into a new commit with a new parent, so the copy's SHA differs from the original. After a few backports, a release branch and the main branch contain the same fixes under different ids. Asking "is this commit upstream?" by id therefore always answers no, which is useless for maintenance work where the real question is "which of my fixes have not been carried across yet". ## The patch-id Git's answer is to hash the **change** rather than the commit. `git patch-id` reads a diff and produces a hash computed after normalising it: hunk headers and line numbers are ignored, and whitespace can be ignored too, so the same edit applied at a different offset in the file still hashes the same. Two commits whose diffs are equivalent in this sense share a patch-id even though their commit SHAs, parents, messages and dates differ. It is a **content heuristic**, not a proof of provenance. Two independently written identical fixes match; a picked commit that was tweaked during the pick — a conflict resolved differently, a variable renamed — does not. ## The porcelain that uses it `git cherry <upstream> [<head> [<limit>]]` is the purpose-built command. It lists the commits in `<head>` that are not in `<upstream>`, prefixing each with: - `+` — no equivalent patch upstream; this one still needs carrying across. - `-` — an equivalent patch is already upstream. `git cherry -v` adds the subject line, which makes the output readable. This is the classic "what still needs backporting" report between a maintenance branch and the main line, in either direction. `git log` exposes the same idea for a symmetric difference between two branches (`git log --oneline A...B`): - `--cherry-mark` marks commits with `=` when an equivalent exists on the other side and `+` when it does not. - `--cherry-pick` omits equivalent commits entirely, so you see only what is genuinely one-sided. - `--cherry` is a shorthand combining the marking with a one-sided, no-merges view. The same detection runs during rebase: when replaying commits, Git skips ones whose change is already present upstream, which is why a rebase of a branch that was partly merged elsewhere can end up shorter than you expected — and why `git rebase` sometimes reports that a commit became empty. ## How to use it in practice Before a release, run `git cherry -v <release-branch> <main-branch>` (or the reverse) to see which fixes exist on one side and not the other. Combine with `-x` provenance lines in backport messages: the two mechanisms are complementary, one recording intent and the other detecting content. `-x` gives an exact, greppable answer for commits you picked deliberately; patch-id gives an approximate answer even for commits nobody annotated. ## Limits worth stating - **Modified picks do not match.** Any adjustment during the pick — a conflict resolved differently, an import fixed up — changes the diff and thus the patch-id. - **Split or squashed changes do not match.** One commit upstream against two here hashes differently. - **Merges are excluded** from these comparisons, since a merge has no single diff. - **Coincidental matches happen.** Two commits that both delete the same trivial line hash the same, so `-` is evidence rather than certainty. - The result depends on the range you ask about; it compares the commits in your specified range, not the whole repository. The interview-worthy point is the conceptual one: Git can compare commits by **identity** (SHA, ancestry) or by **effect** (patch-id), and cherry-pick is precisely the operation that breaks the first while preserving the second. Knowing which commands use which comparison is what lets you answer "is this fix already there" honestly.
- What exactly does a patch-id ignore, and why does that matter?It hashes the diff after normalising away hunk headers and line offsets, and it can ignore whitespace, so the same edit at a different position in the file hashes the same. That is what makes it usable across branches where surrounding code has shifted — which is exactly the situation a backport creates.
- When does this detection give the wrong answer?Whenever the change was altered on the way across: a conflict resolved differently, a rename applied, or one commit split into two. Those produce different patch-ids and show as not-upstream even though the fix is there. It can also match two independently written identical changes, so treat the result as strong evidence, not proof.
- How does this relate to rebase dropping commits?Rebase uses the same equivalence check while replaying: a commit whose change is already present upstream is skipped rather than re-applied. That is why rebasing a branch whose commits were partly cherry-picked into the target can produce fewer commits than you started with, and why Git may report that a commit became empty.
saying these in an interview costs you the question
- Trying to match cherry-picked commits by SHA
- Thinking a patch-id identifies the commit rather than its diff
- Believing the detection is exact rather than a heuristic
- Expecting merges to be compared this way
- Assuming a modified backport will still be detected as present