How does a centralized version-control model differ from a distributed one, and what does centralization still do better?
answer
- where does the history physically live
- which operations need the server
- numbering versus content-derived identity
- files no algorithm can merge
basics
~20 sA centralized model keeps the one history on a server and hands each developer a working copy of a single revision; a distributed model gives every copy the whole history. Centralization still wins at exclusive locking and a single global ordering.
solid answer
~50 sIn a centralized model the server owns the history and each developer holds a working copy of one revision, so reading history, creating a line of development and recording a change are all server round trips. In a distributed model every copy holds the full history, so those operations are local and publishing becomes a separate act. Distribution buys offline work, resilience by replication and near-free lines of development, but it gives up three things a central authority provides naturally: an **exclusive lock** on a file no merge algorithm can combine, a **single total ordering** expressed as a simple increasing revision number, and **path-level access control**, since you cannot hand out part of a history and still let the recipient verify it. Teams that need those add them back as services rather than change models.
go deeper
Be ready to say where the history lives in each model — one server versus every copy — and to name one everyday operation that needs a network in one model and not in the other.
Explain what follows from that: content-derived identity instead of a centrally assigned number, local lines of development, and publishing as a separate act. Expect to be asked what centralization still does better.
Demonstrate judgement about unmergeable assets and access control: a lock needs an authority at the moment of editing, and path-level permissions do not survive a model that hands out whole verifiable histories.
Own the organisational tradeoff — whether to compensate with a locking service, an asset store outside version control, or a repository split — and be explicit about what each choice costs the teams that have to live inside it.
## The one question that separates the models Both models answer the same question differently: *where does the history live?* A centralized version-control system keeps the single history on a server; what a developer holds is a **working copy** — the files as of one chosen revision plus enough bookkeeping to ask the server about the rest. A distributed version-control system gives every copy the entire object store and its pointers, making each copy a peer that can answer historical questions on its own. Almost every practical difference falls out of that one choice. | Property | Centralized model | Distributed model | |---|---|---| | Where the history lives | one server | every copy | | Everyday history operations | server round trip | local | | Identity of a recorded change | an increasing number assigned centrally | a hash derived from content and ancestry | | Ordering across the project | total, by construction | partial; a shared line defines the official order | | Creating a line of development | central and visible immediately | local and private until published | | Recording versus sharing | one act | two acts | | Access control granularity | per directory or path, enforced centrally | per repository in practice | | Unmergeable binary files | an exclusive lock is natural | needs an added locking service | | Server unreachable | most work stops | only exchange stops | | Server lost | history gone unless backed up | any full copy can reseed it | ## What centralization genuinely offers It is a weak answer to say the centralized model has no advantages. It has three, and each is a direct consequence of having one authority that observes every operation. 1. **Exclusive locking for files that cannot be merged.** No algorithm can combine two independent edits to a compiled design asset, an audio file or a packed binary. The only real answer is to stop the second edit from starting. A central server can enforce *reserve this path before editing* because it sees every request. A distributed system has no authority at the moment of editing, so a lock has to be built on top of the shared copy and only binds people who cooperate with it. On a charity donation platform whose campaign artwork bundle runs to 217 MB in a format with no meaningful merge, this is not a theoretical concern — it decides whether two designers can work the same afternoon. 2. **A single, total ordering.** A centrally assigned number is small, speakable and comparable at a glance: revision 4,182 is unambiguously later than 4,181. Content-derived identity is globally unique and verifiable, but it carries no order; ordering exists only along a particular line of history, and the *official* line is a convention the team maintains. Teams recover human-friendly ordering by attaching names to chosen points in history rather than by numbering every change. 3. **Fine-grained access control.** A central server evaluates permissions per path on every operation, so *this team may read only this subtree* is straightforward. Handing out a whole verifiable history is all-or-nothing by construction, so the distributed answer to the same requirement is to split the material into separate repositories — an organisational change, not a setting. ## Why distribution won anyway - **Speed and availability.** The operations engineers perform dozens of times an hour became local, and the shared copy stopped being on the critical path of thinking. - **Lines of development became almost free**, which is what made short-lived, parallel work normal rather than an event. - **Recording and sharing separated**, so unfinished work can be captured safely without being imposed on anyone. - **Replication as a side effect**, which turned the loss of a server from a catastrophe into an inconvenience. - **Contribution without trust.** Someone with no write access to the shared copy can still hold a full history, do complete work in it, and offer the result. ## How teams actually resolve the gap The realistic modern question is not which model to adopt but how to compensate for what distribution removed. - Add an **advisory or enforced locking service** in front of the shared copy for unmergeable assets, and make it a team rule that the tool cannot enforce on its own. - Keep large binary assets in a **store designed for exclusivity and size**, and keep only pointers to them alongside the source. - **Split repositories** where access must genuinely differ, accepting the coordination cost that follows. - Establish an explicit convention for **which line is the official sequence**, since ordering is no longer a property the software supplies. A candidate who can name what was given up, and what the team put in its place, is describing engineering judgement rather than reciting a preference.
- Why is a single increasing revision number hard to reproduce in a distributed system?A number implies a total order agreed by one authority. With independent copies recording changes concurrently there is no authority at the moment of recording, so identity is derived from content instead. Any human-friendly ordering is imposed afterwards by whichever line of history the team treats as official, which makes it a convention rather than a property of the stored objects.
- How do teams with large unmergeable assets cope in a distributed system?They add the missing property rather than change models: a locking service in front of the shared copy, a rule that one named person owns a given asset for a period, or moving those assets into a store built for exclusivity and size while the source history stays distributed. All three work by convention plus tooling, because the model itself offers no lock.
saying these in an interview costs you the question
- Claims the centralized model has no advantages whatsoever
- Thinks distributed means there is no shared server
- Treats a revision number and a content hash as interchangeable identifiers
- Says locking is impossible rather than not built in
- Assumes per-directory permissions behave identically in both models