skip to content

A chat gateway's cluster API answers nothing for an hour while user sessions stay up — which operations stop and which keep working?

level: seniorimportance: must knowfreq 68%

answer

  1. frozen, not down
  2. existing traffic is untouched
  3. every change needs the deciding half
  4. losses during the hour go unrepaired
  5. reading status uses the same API

basics

~20 s

Traffic keeps flowing and running copies keep serving: the node agents and containers need no new decisions. What stops is every change — applying a spec, scaling, placing new work, replacing a copy lost with its host — and reading status, which goes through the same API.

solid answer

~40 s

Nothing in the request path of a running chat session touches the deciding half, so the sessions never notice: the containers are up, their host's node agent is supervising them, and packets reach them by a path already programmed. What stops is everything that *changes* the cluster — applying a changed spec, raising or lowering a copy count, placing anything newly declared, replacing copies lost with a failed host, and advancing a rollout that is mid-flight. You also lose your view: status is read back through the same API. The right phrase is **frozen, not down**. The danger is that the freeze is silent and cumulative: every failure during the hour goes unrepaired, so capacity erodes with nothing replacing it, and a traffic spike arrives with no scaling to absorb it.

code

pseudocode · 11 lines
pseudocode
# an hour with the cluster API unreachable
if request is "serve traffic to an already-running copy":
    continues        # data plane only; the deciding half was never in the path
elif request is "restart a copy that exited on its own host":
    continues        # the node agent enforces the assignment it already holds
elif request is "place a newly declared copy on some host":
    blocked          # placement is a decision, and it must be recorded
elif request is "replace copies lost with a failed host":
    blocked          # same decision, and their agent is gone with the host
elif request is "change the declared copy count" or "read live status":
    blocked          # both are reads or writes through the API

go deeper

for a junior

The headline to remember: an unreachable cluster API does not stop containers that are already running, and users of those workloads usually see nothing at all.

for a middle

Explain why — packets and supervision are local to each host, while placement, scaling and spec changes all require the deciding half and its record.

for a senior

Show the accumulating cost: unrepaired losses, no elasticity, no rollback, a blind dashboard, and a stalled rollout leaving two versions serving at once.

for a principal

Decide in advance how long the estate must tolerate no decisions, and what that buys or costs: spare capacity placed ahead of time, workloads kept off the cluster API's request path, and a rehearsed answer for hour two.

## What the hour actually looks like At minute zero the cluster API stops answering. Every chat session already established stays up. New chat connections still land, because they are carried by copies that were already running and by a network path already programmed on each host. Nobody on the product side notices anything. Meanwhile the operations team can do essentially nothing. They cannot apply a change, cannot scale, cannot read a workload's live status, and cannot see whether the fleet is still the size they think it is. ## What keeps working - **Established and new traffic to running copies.** The deciding half was never in the path of a packet. - **Supervision on each host.** A container that exits is restarted by that host's node agent from the assignment it already holds. - **Local health handling.** Checks configured earlier keep running, and the agent acts on them within the assignment it has. - **Anything the workload does for itself** — its own retries, its own connections to a database, its own timers. ## What stops - **Applying a changed spec.** The write never lands. - **Scaling in either direction.** The declared count lives behind the API. - **Placing anything new.** Placement is a decision, and the decider is unavailable. - **Replacing copies lost with a host.** Their host is gone; re-placing them is a new decision. - **A rollout already in flight.** It stalls where it is, frequently with two versions serving side by side. - **Reading status.** The same API serves reads, so you are blind as well as frozen. | Operation | Which half is needed | During the hour | |---|---|---| | Serve an established session | Data plane only | Continues | | Restart a container that exited in place | Data plane only | Continues | | Read a workload's live status | Control plane | Blocked | | Apply a changed spec | Control plane | Blocked | | Scale up under load | Control plane | Blocked | | Place a newly declared copy | Control plane | Blocked | | Replace copies lost with a failed host | Control plane | Blocked | ## Frozen, not down — and why frozen still costs The honest summary of the hour is that the cluster stopped *changing*, not that it stopped *working*. That distinction is the point of the question, but a strong answer does not stop at the reassurance. Three costs accumulate while nothing appears to be wrong: 1. **Unrepaired loss.** Every copy lost to a crashed host during the hour stays lost. The workload keeps serving on whatever remains, so the erosion is invisible from the outside until the remaining copies are carrying more than they can. 2. **No elasticity.** A spike that would normally be absorbed by scaling is absorbed by queueing and latency instead. 3. **No escape hatch.** If something in the application starts misbehaving at minute twenty, the usual response — change the spec, roll back, scale out — is unavailable for the rest of the hour. So the risk profile is the opposite of what it looks like: the first ten minutes are genuinely harmless, and the sixtieth minute is not. ## What you cannot see Being blind deserves separate billing. Status flows from the node agents, through the API, into the state store, and back out to whoever asks. With the API unavailable, the whole chain is unavailable, and any dashboard reading it freezes at its last value. Teams that route workload metrics and logs through a collector independent of the cluster API keep their view during exactly this hour — which is the practical reason those paths are usually kept separate. ## Designing so the hour stays boring - **Keep enough copies, already placed, already spread.** Capacity that already exists needs no decision. - **Keep the cluster API out of the request path.** A workload that queries the deciding half while serving a user request converts a control-plane outage directly into a user-visible one. The chat gateway in this scenario survived because it does not do that. - **Pre-scale ahead of a known peak** rather than relying on reacting to it. - **Practise the hour.** The question interviewers ask next is *how do you know your fleet survives it*, and the only convincing answer is that you have tried. ## The trap in the wording The phrase people reach for is *the cluster is down*. It is wrong and it leads to wrong actions — restarting hosts, draining traffic, failing over a product that is still serving perfectly. Say *the control plane is unavailable; the data plane is unaffected*, then list what that costs.

  • Why is a four-hour outage of the deciding half far worse than a four-minute one?
    Because risk accumulates rather than arriving at once. Each copy lost to a crashed host during the window stays lost, so capacity erodes with no replacement and no signal, load concentrates on what remains, and there is no lever to scale out or roll back when the first of those tips over.
  • What design choice would have made the chat gateway notice the hour immediately?
    Calling the cluster API on the request path — for configuration, discovery or a permission check while serving a user. That turns an outage of the deciding half into a user-visible outage. Workloads that read what they need at start-up, or from a source outside the cluster API, keep serving through it.
  • Does a rollout that was half finished continue during the hour?
    No. Replacing the next batch is a decision, so the rollout stalls exactly where it was, commonly with the old and new versions both serving. It resumes when the API returns — which is why a mixed-version window is worth designing for rather than assuming it lasts only minutes.

A control tower going silent does not bring down the aircraft already in the air. They keep flying on the clearances they were given; what stops is anything taking off, and anyone being given a new heading.

saying these in an interview costs you the question

  • Says the cluster is down, so the service must be down.
  • Expects autoscaling to absorb a traffic spike during the outage.
  • Thinks copies lost with a failed host come back on their own.
  • Assumes live status can still be read while the API is silent.
  • Treats the hour as harmless because nothing broke immediately.
  • Believes an in-flight rollout finishes itself and then stops.