Why is every clone in a distributed version-control system a full repository, and what does that buy you?
answer
- the copy is not a subset
- history travels with the files
- which operations need a network
- offline, resilience, cheap experiments, replication
basics
~20 sA clone copies the entire history and every stored object, not just the newest files, so the copy is a working repository on its own — you can record changes, read history and switch lines of development with no server.
solid answer
~50 sCopying a repository in a distributed system transfers the whole object store plus the pointers that name positions in it, so the new copy holds the same history as the original rather than a slice of it. Four things follow. **You can work offline**: recording a change, reading history and comparing two old revisions are local operations. **There is no single point of failure**: every complete copy holds the history it has received and could take over as the shared one. **Experiments are cheap**: creating and discarding a line of development writes one small pointer on your own disk and is invisible to everyone else. **Publishing is explicit**: recording a change and sharing it are separate acts. The price is disk space and a first transfer sized by the whole history, which only bites on very old or asset-heavy repositories.
go deeper
Be ready to say that a clone brings the whole history and not just the current files, and to name two everyday operations that therefore need no network at all.
Explain the mechanics: an object store plus movable pointers, both copied, and which operations become local as a result. Expect a follow-up on what completeness costs in disk and first-copy time.
Show the operational side — what a team can still do while the shared copy is unreachable, and why replication across developer machines is redundancy rather than a backup with retention and verified restores.
Own the tradeoff at scale: whole-history copies stop being free on very old or asset-heavy repositories, and the answer is a decision about repository size, splitting, and what history the organisation is obliged to keep.
## What a clone actually contains A repository in a distributed version-control system is two things kept side by side. The first is an **object store**: an append-only collection of immutable records holding every version of every file the project has ever contained, plus the records that chain those versions into a history. Objects are addressed by a hash derived from their own content, so identical content always lands in the same place and a copy can be checked for integrity without asking anyone. The second is a set of **pointers**: short, movable names — a line of development, a name for a fixed point in history — that record where in that history something currently sits. Copying the repository transfers both. That is the structural difference from a centralized model, where what arrives on your machine is a **working copy**: the files as of one chosen revision, plus enough bookkeeping to ask the server about everything else. In a distributed system there is no everything-else living somewhere you cannot reach. Your copy is a peer of the copy you took it from, identical in kind, differing only in where its pointers sit and which recorded changes it holds that the other does not. ## The four things completeness buys 1. **Offline work.** Reading history, comparing two old revisions, finding when a line of a file changed, creating a line of development and recording a change are all reads and writes against local storage. A network is needed only to exchange work with another copy. 2. **No single point of failure.** History is replicated as a side effect of working. If the machine hosting the shared copy is lost, any complete copy still holds the history up to what it last received and can become the new shared copy. 3. **Cheap local experimentation.** Creating a line of development writes one small pointer, and nothing you record is visible to anyone until you publish it, so trying an idea and throwing it away costs only your own disk. There is no shared namespace to pollute and nothing to clean up in public. 4. **Publishing is a deliberate act.** Recording a change and sharing it are separated, which is the same property seen from the other side: your copy diverges on purpose and you choose when it converges again. ## Where the cost lands | Operation | Centralized working copy | Distributed full copy | |---|---|---| | Read the history of a file | server round trip | local | | Compare two old revisions | server round trip | local | | Create a line of development | created centrally, visible to all | local, private until published | | Record a change | enters the shared sequence at once | local; publishing is a separate act | | First copy of the project | one revision's worth of files | every version ever recorded | | Shared copy unreachable | most work stops | only exchange stops | | Shared copy lost | restore from backups | any full copy can reseed it | The last two rows are the payoff and the first-copy row is the price. On a charity donation platform whose shared copy was unreachable for 43 minutes during a seasonal traffic peak, no engineering time was lost: 12 people kept recording changes, reading history and preparing work, and the exchange queue drained when the host returned. The same repository's first copy moves the whole history, including versions of files deleted years earlier, which is why the cost of completeness shows up on old, large or asset-heavy projects and almost nowhere else. ## What completeness does not give you Being complete is a statement about the past, not about the present. Your copy holds everything it has received; it says nothing about what a teammate recorded ten minutes ago and has not published. Two consequences are easy to miss. - **Replication is redundancy, not a backup policy.** Developer copies drift, they never contain unpublished work, and they faithfully replicate a destructive rewrite the moment they receive it. Redundancy with no retention, no verification and no restore drill is not a backup. - **Completeness coordinates nobody.** Holding the whole history does not tell you which copy the team treats as authoritative, whether a change has been accepted, or whether someone else is already editing the file you just opened. Those are conventions and services layered on top of the model. A 14-month-old line of development on that donation platform shows both at once: every copy that ever received it still holds it in full, which is why nothing was lost — and no copy can say whether the team still wants it, which is why nothing was decided.
- What does a full copy cost you that a single-revision checkout does not?Disk space and a first transfer sized by the entire history, including versions of files that were deleted long ago. After that, exchanges are incremental and small. The cost is therefore concentrated in the first copy and grows with age and with binary assets, not with the number of people working.
- Does having a complete copy on every machine count as a backup strategy?Only as redundancy. Copies drift, they hold nothing that was never published, and they replicate a destructive rewrite as readily as a good change. A backup strategy needs retention, a verified restore and a copy that is not updated by the same act that damaged the original. Replication gives none of those.
A centralized checkout is a library reading room where you borrow one book at a time. A clone is being handed the whole library, catalogue included.
saying these in an interview costs you the question
- Says a clone downloads only the newest version of the files
- Believes reading history always requires contacting a server
- Thinks distributed means no shared copy exists at all
- Claims every developer's machine is a verified backup
- Assumes recording a change locally makes it visible to teammates