When two copies of a distributed version-control repository synchronise, what actually moves between them?
answer
- two steps, not one
- objects first, then names
- your record of their positions
- integration is a separate local decision
basics
~20 sSynchronising transfers the history objects one side lacks and then updates pointers. Retrieving from a shared copy moves your record of where their lines of development sit; your own line and your working files change only when you integrate.
solid answer
~50 sA synchronisation is two steps, and keeping them apart is the whole point. First an **object exchange**: the sides work out which history objects one holds and the other lacks, and transfer only those. Second a **pointer update**: the receiving side moves the short names that record where a line of development sits. When you retrieve from a shared copy, the pointers that move are your local record of where their lines are — your own line of development and your working files are untouched, which is why retrieving can never confront you with a decision. Integrating is the separate, local act that follows: you join their history to yours, by merging the two lines of history or by replaying your changes on top. Publishing runs the same two steps outward, with one extra rule — the receiver normally refuses a pointer move that would drop recorded history.
code
pseudocode · 12 linesbefore retrieving
my line of development -> C7 (C5, C6 not yet seen)
my record of shared line -> C4
after retrieving (objects C5 and C6 arrive)
my line of development -> C7 unchanged
my record of shared line -> C6 moved
working files unchanged
after integrating
my line of development -> C8 joins C7 and C6
working files updatedgo deeper
Be ready to say that synchronising moves history objects and then updates names, and that retrieving another copy's work does not change the files you are currently editing.
Explain the two steps separately: how the sides decide which objects are missing, which pointers move on the receiving side, and why integrating retrieved history is a further local operation that can require a decision.
Show why the split is operationally useful — transfers are safe to schedule and to retry, an interrupted exchange leaves spare objects rather than damage, and ahead-behind counts reflect your last refresh rather than live state.
Own the coordination question: when a shared pointer may be moved off its own history at all, who is allowed to decide that, and what it costs every copy that already received the history being abandoned.
## Synchronisation is two steps, never one When two repositories synchronise, exactly two kinds of thing happen, in this order. 1. **An object exchange.** The two sides negotiate which history objects one holds and the other does not, and transfer only the missing ones. Because objects are immutable and named by a hash of their own content, *missing* is a well-defined question with a cheap answer, and an object that arrives twice is stored once. 2. **A pointer update.** Once the objects are safely stored, the receiving side moves the short names that record where each line of development sits. A pointer move is a tiny write, and it is the only part of the operation that changes what anything means. Everything people say about retrieving, integrating and publishing is a statement about which side performs which step, and which pointers move. ## Retrieving versus integrating When you retrieve from another copy, the objects land in your store and the pointers that move are **your local record of where their lines of development are** — not your own line, and not your working files. That is why retrieval is safe: afterwards your copy knows more and behaves no differently. You can read what arrived, compare it with your own work, and decide at leisure. Integrating is the separate, local decision that follows. You join their history to yours — either by merging two lines of history into a new recorded change, or by replaying your own changes on top of theirs — and only now can two people's edits to the same region collide, and only now do your working files change. The two are commonly offered together as one convenience operation, and that is where surprises come from. The transfer half can never require a decision; the integration half can, and it moves your own line of development. Anyone who says that retrieving changed their files is describing the second half. | Act | What transfers | Which pointers move | Can it require a human decision? | |---|---|---|---| | Retrieving | objects you lack | your record of their positions | No | | Integrating | nothing over the network | your own line of development | Yes | | Publishing | objects they lack | a shared pointer on the other side | Yes, if the move is refused | ## Publishing is the same exchange with one rule Publishing sends the objects the other side lacks and then asks it to move a shared pointer. The receiver normally accepts that move only when the new position already contains the pointer's current position — in other words, only when the move extends the existing history rather than stepping sideways off it. If it stepped sideways, recorded changes that other copies already hold would stop being reachable from that name, and every copy that already received them would keep trying to put them back. That refusal is a **history-safety rule, not a permissions check**. Someone with full write access still meets it, and the fix is to bring the other side's work into yours first and publish the combined result. Overriding the rule is possible and is a coordination decision rather than a technical one, because it silently invalidates what other copies believe. ## Reading ahead-and-behind counts A report that you are ahead by 3 and behind by 11 is computed against **your last retrieved record** of the other copy, not against that copy live. On a charity donation platform, an engineer who had not retrieved for eleven days read *behind by 4* while the shared line had in fact moved on by 1,284 recorded changes during a seasonal traffic peak. Nothing was broken; the number was faithfully answering a question about stale local data. Refresh the record first, then read the comparison. ## Why the split matters operationally - **Transfers are safe to automate.** A machine that only exchanges objects never lands in a state a human must repair, so retrieving on a schedule is cheap and boring. - **You can inspect before you adopt.** Incoming work sits in your object store where you can read and compare it without it touching your files, which is the basis of examining someone's change locally. - **Interrupted exchanges are recoverable.** Objects are written before any pointer moves, so a broken link leaves you with extra unreferenced objects rather than a damaged history, and the next attempt transfers only what is still missing. - **Bandwidth follows novelty, not size.** Because both sides compare what they already hold, a second synchronisation of a large repository moves almost nothing.
- Why would a receiving repository refuse a publish even from someone with full write access?Because the requested pointer move would step off the pointer's own history and make recorded changes unreachable from that name. Every copy that already received those changes still holds them, so the refusal protects shared work rather than enforcing permissions. The normal fix is to bring in the other side's work and publish the combined result.
- If retrieving never touches your work, why do people still say it broke their files?Because retrieving and integrating are usually offered as a single convenience operation, so the integration half runs unnoticed. That half moves your line of development, rewrites your working files and can stop on a collision. Splitting the two — transfer, look, then decide — makes the surprise impossible.
Retrieving is the post being delivered to your hallway; integrating is opening the letters and acting on them. Only the second one can ruin your afternoon.
saying these in an interview costs you the question
- Thinks retrieving from a shared copy rewrites your working files
- Says every synchronisation copies the whole history again
- Cannot separate moving a pointer from transferring objects
- Believes a publish always succeeds if you have write access
- Assumes behind-by-two is measured live against the shared copy