skip to content

How would you maintain a long-lived fork carrying local patches on top of an actively developed upstream project?

level: principalimportance: should knowfreq 30%

answer

  1. A fork is a liability with a running cost
  2. Separate what upstream owns from what you own
  3. Keep local change enumerable, not scattered
  4. Replay the series onto each new release
  5. git rebase --onto plus rerere.enabled

basics

~20 s

Keep local changes as a small, curated commit series on a branch above an untouched mirror of upstream, refresh it by rebasing onto each new upstream release, measure the recurring conflict cost, and continuously shrink the delta by getting patches accepted upstream.

solid answer

~50 s

Treat the fork as upstream plus a patch series, never as a repository of its own. Mirror `upstream/main` and the release tags untouched, and hold every local change on a branch above that base, with each commit small, single-purpose and self-describing. Refresh by fetching upstream and running `git rebase --onto upstream/<new-tag> <old-base> <patch-branch>`, which replays the series onto the new release; enable `rerere.enabled` so recurring conflict resolutions get replayed automatically. The metric I actually manage is the size of the delta — patch count and how much conflict work each upstream bump costs — and the strategy is to drive it toward zero by upstreaming what the project will take and reimplementing the rest as configuration or extension points rather than edits to upstream files. If the delta must stay large and auditable, merging upstream in each cycle is the alternative: cheaper per cycle, much worse to read over years.

code

bash · 2 lines
bash
git fetch upstream --tags
git rebase --onto upstream/v3.5.0 upstream/v3.4.0 patches

go deeper

for a junior

Know that a fork does not update itself and that local changes must be replayed onto newer upstream code. Recognise that keeping many local changes over a long time is expensive.

for a middle

Explain the mechanics of replaying a series onto a new base with git rebase --onto, why the old base has to be named, and why conflicts recur per patch across refreshes.

for a senior

Show you can operate this: mirror untouched, curated patch series, scheduled refresh onto tags, rerere for repeated conflicts, and a diagnosis when a refresh starts costing more each cycle.

for a principal

Own it as a standing liability with a budget. Argue about delta size as the managed metric, upstreaming as the only permanent fix, configuration over modification, and a written exit strategy with a named owner.

## Why long-lived forks fail A long-lived fork fails slowly. Each upstream release costs a little more to absorb than the last, until someone decides that skipping this release is fine — and then you are stuck on an old version with no cheap way forward. Every design choice below exists to keep that cost visible and bounded. ## Model the fork as base plus patch series The discipline that makes everything else possible: your fork's `main` is a **mirror**, byte-identical to upstream, never committed to. All local change lives on a separate branch as an ordered series of commits above a named upstream base: ``` upstream/v3.4.0 <- base |- fix: disable telemetry by default |- feat: add internal auth provider |- chore: pin build toolchain ``` Each commit is small, single-purpose, and its message says *why* it exists locally and what would make it unnecessary. That last part matters: a patch whose exit condition is written down is a patch someone can retire. ## The refresh operation When upstream releases: ``` git fetch upstream --tags git rebase --onto upstream/v3.5.0 upstream/v3.4.0 patches ``` `git rebase --onto` is the right tool because it names the old base explicitly: replay everything after `upstream/v3.4.0` on top of `upstream/v3.5.0`. Conflicts are resolved per patch, which localises each break to the specific local change that no longer applies — exactly the signal you want. Turn on `rerere.enabled` so Git records each conflict resolution and reapplies it when the same conflict reappears in the next cycle. On a series that gets rebased repeatedly, that saves real work. ## The alternative: merge instead of rebase Merging each upstream release into a permanently divergent branch is the other model. It resolves conflicts once per cycle instead of once per patch, never rewrites history, and produces an audit trail of exactly when each upstream version was absorbed — which is sometimes a regulatory requirement. The price: after a few years nobody can answer "what have we actually changed?" without a diff against the upstream tag, because the local changes are no longer a readable series. Choose merging when auditability and low per-cycle friction outrank the ability to enumerate your delta; choose rebasing when the delta is the thing you must keep understandable. ## Manage the delta as the real metric The number to watch is not lines changed but **how many patches you carry and how much each upstream bump costs**. Track it deliberately and act on it: - **Upstream everything you can.** An accepted patch is the only permanent removal from your maintenance burden. - **Prefer configuration to modification.** A change expressed through an existing extension point, plugin interface, or config file survives upstream refactors untouched; a change edited into an upstream source file conflicts every time that file moves. - **Prefer additive files to edits.** New files you own rarely conflict; edited upstream files always can. - **Retire patches.** When upstream implements your feature, delete yours instead of keeping both. ## Pin to releases, not to a moving tip Rebase onto tagged releases rather than upstream's tip. Tags give you a stable, reproducible base, a defined cadence, and a name you can put in a changelog. Following the tip means absorbing every intermediate breakage and never having a clean statement of what you are based on. ## Make the cost visible to the organisation The fork is a standing liability, so treat it like one: schedule the refresh rather than doing it under deadline pressure, keep the patch series and its exit conditions where reviewers can read them, and have a named owner. Automate the mechanical part — a scheduled job that attempts the rebase onto the newest tag and reports failures — so drift is discovered by a robot on a Tuesday, not by an engineer during an urgent security update. ## Know the exit strategies A long-lived fork should always have a written answer to "how does this end": all patches upstreamed and the fork deleted; the local changes extracted into a plugin or wrapper that depends on upstream normally; a hard fork with a new name, accepting that you now own the whole codebase; or a migration off the dependency. A fork with no exit strategy is one that will be abandoned mid-version. ## What interviewers want This is a judgment question with no single right answer. Strong answers show a mechanism (mirror plus patch series, `rebase --onto` per release, `rerere` for repeated conflicts), a metric (delta size and refresh cost), a strategy for shrinking it (upstream, configure, add rather than edit), and an exit plan. Weak answers describe a one-off sync command and never address what happens on the tenth upstream release.

  • Why rebase onto upstream tags rather than continuously onto upstream's tip?
    Tags give a reproducible, nameable base with a defined cadence, so you can state exactly what your fork is built on and absorb one reviewed set of changes at a time. Tracking the tip means re-resolving conflicts against intermediate states, including ones upstream later reverted, and you never have a stable answer to what version you are running.
  • When would you accept merging upstream into a divergent branch instead of maintaining a rebased series?
    When auditability of when each upstream version was absorbed matters more than being able to enumerate your changes, when the delta is too large to rebase economically, or when several people work on the fork concurrently and rewriting the branch would disrupt them. You trade a readable patch series for one conflict resolution per cycle.
  • What signals tell you a long-lived fork should be ended rather than maintained?
    Rising conflict cost per upstream release, patches whose original justification nobody can state, skipping releases to avoid the refresh, and security updates blocked behind a rebase. At that point the choices are to upstream aggressively, re-express the changes as a plugin or wrapper on stock upstream, or commit to a hard fork with full ownership.

saying these in an interview costs you the question

  • Committing local changes directly onto the fork's mirrored main
  • Never upstreaming, so the delta only grows
  • Chasing upstream's tip instead of tagged releases
  • Treating conflict pain as unavoidable rather than as a metric
  • Maintaining a fork with no written exit strategy

context