skip to content

Three-phase commit (3PC) adds a 'pre-commit' phase between the initial prepare vote and the final commit, specifically to let surviving participants recover without waiting indefinitely for a crashed coordinator. Explain how that extra phase helps under a simple crash-stop failure model, and why 3PC still doesn't solve the blocking problem once you allow network partitions.

level: seniorimportance: should knowfreq 35%

answer

  1. pre-commit phase = extra ack round
  2. non-blocking only under crash-stop
  3. partitions cause split-brain: some commit, some abort
  4. assumes reliable non-partitioned network
  5. rarely used in real production systems

basics

~20 s

3PC adds a middle step so participants know 'everyone agreed to commit' before actually committing, letting them safely finish on their own if the coordinator dies. But if the network splits instead of the coordinator crashing, participants on different sides can still disagree, so it still gets stuck.

solid answer

~50 s

3PC splits 2PC's commit phase into pre-commit and commit, so a participant only needs to know whether any participant reached pre-commit — not the coordinator's final word — to safely resolve a coordinator crash: if it or a peer already reached pre-commit, it's safe to commit (no one could have aborted after that point); if no one did, it's safe to abort. Under a clean crash-stop model (only the coordinator fails, network is reliable), this bounds the blocking window because a timeout-driven election among survivors can always determine the safe action. But under network partitions, participants can split into groups that each believe they're the surviving majority — one side times out and aborts while the other times out and commits — producing exactly the split-brain inconsistency 3PC was meant to avoid. That's why 3PC is a textbook protocol rarely used in production networks, where partitions are common.

go deeper

for a junior

Should know 3PC adds an extra step compared to 2PC and that it's meant to reduce blocking, without needing the partition nuance.

for a middle

Should be able to describe the three phases (CanCommit/PreCommit/DoCommit) and give the basic intuition of why pre-commit helps.

for a senior

Should explain precisely why pre-commit makes local recovery safe under crash-stop, and construct the partition scenario that breaks it (split-brain commit vs abort).

for a principal

Should connect this to consensus theory — why partitions require quorum-based agreement, not just extra communication rounds — and explain why production systems route around 3PC entirely with consensus-backed coordinators or Sagas.

## Why 2PC blocks Two-phase commit blocks because a participant that has voted COMMIT has no way to distinguish 'the coordinator decided COMMIT and crashed before telling me' from 'the coordinator decided ABORT and crashed before telling me' — both are consistent with everything the participant has observed, and guessing wrong breaks atomicity. Three-phase commit's fix is to insert an intermediate phase, **pre-commit**, whose entire purpose is to make that ambiguity structurally impossible under the failure model 3PC assumes. ## The three phases The protocol now runs: | # | Phase | What happens | |---|---|---| | 1 | **CanCommit** | the coordinator asks for votes, like 2PC's prepare | | 2 | **PreCommit** | if every vote was YES, the coordinator sends PRE-COMMIT to all participants and waits for acknowledgments, and a participant receiving PRE-COMMIT knows for certain every other participant also voted YES, because the coordinator would not have sent it otherwise | | 3 | **DoCommit** | once the coordinator has all pre-commit acks, it sends the final COMMIT | ## What pre-commit guarantees The invariant 3PC engineers in is: - no participant ever reaches pre-commit unless every participant is guaranteed to eventually commit, and - no participant commits unless it (or, transitively, some participant it can learn from) reached pre-commit. So if the coordinator dies, surviving participants can look at their own local state — did I reach pre-commit? — and, by talking to each other, agree on a safe action without needing the coordinator: - if any survivor reached pre-commit, it's always safe for everyone to proceed to commit; - if none did, it's always safe to abort, since nothing has been promised yet. A timeout-based election lets survivors pick a new coordinator, poll each other's states, and resolve the transaction — bounded by the timeout rather than by waiting for the original coordinator to return. This is the sense in which 3PC is 'non-blocking': under the failure model it assumes. ## Where partitions break it That failure model is the crucial caveat: **crash-stop**, with reliable, non-partitioned message delivery — a process either works correctly or halts completely, and every message sent between live processes eventually arrives. Real networks don't behave that way; they partition. Once you allow partitions, 3PC's non-blocking guarantee unravels. Consider a coordinator that sends PRE-COMMIT to participant A, A acknowledges, and then a partition isolates A from the rest of the cluster (coordinator and participant B) before the coordinator hears back from B. On A's side, A has reached pre-commit — it will eventually time out, elect itself or a peer as leader, see it's in pre-commit, and safely commit. Meanwhile, on the other side, B and the coordinator never received a pre-commit acknowledgment from A and can't confirm the state; B's side may time out and decide to abort, believing pre-commit was never reached everywhere. The result: A commits, B aborts — **split-brain**, the exact outcome the protocol exists to prevent. The root cause is that 3PC's safety argument relies on 'a participant hears from enough others to know the true global state,' and a partition can make that assumption false on both sides simultaneously — each side can look locally consistent while the two sides disagree. ## The broader impossibility This reflects the broader impossibility that you cannot get both non-blocking termination and safety in an asynchronous network where partitions (not just process crashes) are possible, without additional machinery like a quorum-based consensus protocol (Paxos, Raft) underneath the coordinator itself. In practice this is why 3PC, despite being the standard textbook answer to '2PC blocks, how do you fix it,' essentially never appears in production distributed databases — real infrastructure either: - accepts 2PC's blocking risk and invests in fast coordinator recovery, - replaces the coordinator with a replicated, consensus-backed log (so partitions are resolved by quorum majority rather than participant guesswork), or - sidesteps the whole problem by using compensable, eventually-consistent Sagas instead of atomic distributed commit.

  • Why is it safe for a participant to commit as soon as it reaches the pre-commit state, without waiting for the final COMMIT message?
    Reaching pre-commit is only possible if the coordinator had already collected a unanimous YES vote from every participant, so by the time any one participant is in pre-commit, it's guaranteed no participant voted ABORT and none can retroactively do so. That guarantee is what lets a surviving participant safely commit on its own during coordinator-crash recovery.
  • What extra infrastructure do real systems use instead of 3PC to avoid coordinator blocking?
    Most replace the single coordinator with a quorum-based consensus protocol like Paxos or Raft, so the 'coordinator role' is itself replicated and survives individual node failures without ambiguity. Google's Percolator and Spanner, for example, build distributed transactions on top of a consensus-replicated timestamp/lock service rather than a single fragile coordinator.
  • Does adding more participant-to-participant communication rounds fix the split-brain scenario in 3PC?
    Not fundamentally — more rounds can reduce the window of risk but cannot eliminate it in an asynchronous, partition-prone network, because a partition can always be timed to fall exactly between whichever rounds were added. Only quorum-based agreement, where only one side of any partition can ever have a majority, actually closes the gap.

Like a stage director who, before the final cue, checks that every actor has silently nodded 'ready' (pre-commit) so that if the director collapses, any actor who saw everyone nod can confidently start the scene alone. But if a curtain (a network partition) drops between two actors before all the nods are compared, each side of the curtain can start the scene independently and out of sync with the other.

saying these in an interview costs you the question

  • Claims 3PC is non-blocking in all real-world conditions
  • Doesn't know what the pre-commit phase actually guarantees
  • Thinks 3PC and 2PC have the same number of message round trips
  • Believes network partitions and coordinator crashes are the same failure mode
  • Can't explain why split-brain (some commit, some abort) is possible under 3PC + partition

context