Describe the lifecycle/states of a mirror topic and what 'promote' versus 'failover' do.
answer
- ACTIVE → STOPPED makes it writable
- Promote = graceful, drains lag, RPO 0
- Failover = emergency, source down, tail loss
- Pause/Resume ≠ writable
- Promote needs healthy source
basics
~20 sA mirror topic starts in an ACTIVE (mirroring, read-only) state. To make it writable you stop mirroring and convert it: 'promote' is a graceful, synced cutover (no data loss); 'failover' is an immediate cutover used when the source is unreachable, accepting possible loss of un-replicated records.
solid answer
~50 sA mirror topic has a mirror state. While mirroring it is ACTIVE and read-only — the destination continuously pulls from the source. To turn it into a normal writable topic you run a state transition. **Promote** is the planned, graceful path: the destination verifies it has caught up with the source (mirror lag drained), then stops mirroring and makes the topic writable — guaranteeing no data loss. **Failover** is the emergency path used when the source cluster is down or unreachable: it immediately stops mirroring and makes the topic writable without waiting to drain lag, so any records not yet replicated are lost. After either transition the mirror state becomes STOPPED, the link no longer pulls into that topic, and applications can produce to it. You can also PAUSE/RESUME mirroring without making the topic writable. Choosing promote vs failover is fundamentally an RPO decision: promote = zero RPO but needs a healthy source; failover = nonzero potential RPO but works when the source is gone.
go deeper
Know mirror topics are read-only until you convert them, and that the convert step is promote/failover.
Explain ACTIVE/PAUSED/STOPPED and the promote-vs-failover safety difference clearly.
Frame promote vs failover as an explicit RPO/availability decision and know the catch-up/lag mechanics.
Design runbooks: when to drill promote vs failover, timeout handling, and re-establishing reverse replication after cutover.
## Mirror state — what it is Every mirror topic carries a **mirror state** that controls whether it is still being fed from the source and whether it is read-only. The important states: - **ACTIVE** (a.k.a. mirroring): the topic is read-only and the destination is continuously pulling records from the bound source topic over the link. New records appear as the source produces them. - **PAUSED**: mirroring is temporarily halted (no new data pulled) but the topic remains read-only and the binding still exists. RESUME returns it to ACTIVE. - **STOPPED**: the mirror binding is broken; the topic is now a **normal, writable topic** on the destination. There is no going back to mirroring for that topic. - **FAILED / PENDING** transitional states may appear during a cutover. ## Promote vs failover — the two ways to STOP into writable Both turn a read-only mirror into a writable topic, but they differ in safety: ### Promote (graceful, planned cutover) 1. The command checks **mirror lag** — how far behind the destination is relative to the source's high-water mark. 2. It waits until lag is drained (destination has fetched everything available from the source). 3. It stops mirroring and makes the topic writable. Because it confirms catch-up first, **promote guarantees no data loss (RPO = 0)**. It requires the **source to be reachable and healthy** — you're coordinating a clean handoff, e.g. during a planned migration or a controlled DR drill. ### Failover (emergency cutover) 1. Used when the **source cluster is unavailable** (region outage, etc.). 2. It **immediately** stops mirroring and makes the topic writable, **without** waiting to drain lag. 3. Any records produced to the source that hadn't yet been replicated are **lost**. So failover trades a possibly **nonzero RPO** for availability — you accept losing the un-replicated tail in exchange for being able to recover when the source is gone. ## The decision: it's an RPO/availability trade-off - **Source healthy, planned cutover** → promote (zero loss). - **Source dead, must recover now** → failover (accept tail loss). ## Edge cases - After STOPPED there's no automatic 'fail back'; re-establishing replication the other direction means setting up a **new link/mirror** in reverse. - Promote can block/timeout if the source keeps producing faster than the destination drains, or if the source is unreachable (then you'd use failover). - PAUSE/RESUME does **not** make the topic writable — only promote/failover do. ## CLI shape (illustrative) With the Confluent CLI you transition mirror topics, e.g. `confluent kafka mirror promote <topic> --link <link>` and `confluent kafka mirror failover <topic> --link <link>`.
- Your source region just went fully offline and you must accept writes on the destination immediately. Promote or failover?Failover — promote requires a reachable source to drain lag, which you don't have. Failover accepts the possible loss of un-replicated records to restore availability.
- What does PAUSE do that promote/failover don't?PAUSE stops pulling new data but keeps the topic read-only and the mirror binding intact, so it can RESUME. Promote/failover permanently stop mirroring and make the topic writable.
saying these in an interview costs you the question
- Saying promote and failover are interchangeable
- Claiming failover guarantees no data loss
- Thinking a paused mirror topic is writable
- Believing you can resume mirroring on a STOPPED (promoted) topic