skip to content

Control Plane Dependency

The management API degrading while already-running workloads keep serving, so scaling, deploying and replacing an instance break first. Asked because a static posture is the non-obvious answer.

on this pageshow

questions

4

During a provider incident, running workloads keep serving but no new instance launches - which half of the platform is degraded, and what stops?

level: middleimportance: must knowfreq 62%

answer

  1. two halves, different blast radius
  2. traffic path versus management path
  3. already running keeps running
  4. create, change, delete are the casualties
  5. control plane degraded, data plane serving

basics

~20 s

The control plane - the platform's management API that creates, changes and deletes resources - is degraded, while the data plane that carries request traffic keeps serving. Launching, scaling, replacing and failing over stop; already-running capacity does not.

solid answer

~50 s

A platform is two systems. The **data plane** carries your traffic: processes already running, disks already attached, packets already routed. The **control plane** is the management API that creates, changes and deletes resources, with a console and a command line as clients of it. In most provider incidents the control plane degrades first and hardest, because it is the coordinated, region-wide, strongly consistent part, while the data plane runs on state each host already holds. So requests keep being served and objects keep being read, but launching a worker, running a rolling deployment, replacing a failed instance, registering new capacity behind an entry point and any failover that has to promote or provision something all stop. The correct sentence in the incident channel is `the control plane is degraded`, not `the region is down` - those imply completely different responses.

go deeper

for a junior

Recall the two names and which one you are usually talking to. The console and command line are clients of the management API; the traffic your users generate does not go through it.

for a middle

Explain the split mechanically: which operations are really create-change-delete calls, why those are the ones that stop, and why already-running capacity is unaffected.

for a senior

Show the operational reflex: diagnose from your own success rate rather than a status page, name what is frozen, and say what you will stop your automation from doing while you cannot create anything.

for a principal

Frame it as a design constraint you impose on other teams - mitigations that change existing state rather than request new state, and headroom bought before the incident rather than during it.

## The two planes Every large cloud platform is built as two systems that fail independently, and the whole of this topic follows from that split. The **data plane** is the part that carries your work: the processes already running on rented machines, the disks already attached to them, the packets already flowing along routes that were installed earlier, the bytes already sitting in an object store and being read. Its work is executed close to where the state lives, by components that already hold the configuration they need. The **control plane** is the management API: the interface through which resources are created, changed and deleted. A console, a command line and an infrastructure tool are all just clients of that same API. Its work is coordinated - a create has to pick a physical host, reserve capacity, allocate addresses, update several inventories and make the result visible consistently across a region. | | Control plane | Data plane | |---|---|---| | What it does | Creates, changes and deletes resources | Carries application traffic and serves stored bytes | | A typical request | Launch a worker; attach a volume; change a routing rule | Serve this request; read this object; deliver this packet | | Where its state lives | Coordinated, often region-wide, strongly consistent | Mostly local to a host, instance or path | | What failure looks like | Errors, long hangs, calls that half-complete | Dropped requests, higher latency, unreachable data | ## What actually stops When only the control plane is degraded, the operations that break are the ones that are secretly management calls: - **Scale-out.** Adding a worker is a create call, so a fleet is frozen at the size it already had. - **Rolling deployment.** Every batch retires instances and creates new ones; a deploy started now can strand itself half-done. - **Instance replacement.** A supervision loop that terminates an unhealthy worker and creates a fresh one performs the destructive half successfully and the constructive half not at all. - **Registering new capacity** behind the entry point in front of a fleet, which is a configuration change on a managed component. - **Configuration a new instance reads at boot**, which frequently comes from a management endpoint rather than from the workload's own data path. - **Failover that provisions or promotes something** - standing up a replacement, promoting a standby, repointing a name - because each of those is a change, not a request. What keeps working is everything that was already in place: instances serving, queues delivering, attached disks reading and writing, existing routes carrying packets, caches answering. ## Why this is the usual shape The asymmetry is deliberate on the provider's side, and it is the same property you are supposed to copy. The data plane is built to be **statically stable**: it keeps doing the last thing it was told to do without needing to ask anything. A host does not consult a central service to keep a process running. A route does not need re-approval to carry the next packet. The control plane cannot be built that way, because its entire job is to change global state consistently, which means coordination, which means a much larger shared failure domain. There is one dangerous asymmetry inside the degradation itself: destructive and idempotent calls often keep landing while creates fail or hang. A fleet can therefore shrink during an incident and be unable to grow back. ## What it means for your design 1. **Assume you can create nothing.** Design every incident response so the mitigation is a change to something that already exists, not a request for something new. 2. **Buy the capacity before the incident.** Pre-provisioned headroom and a warm standby are the only kinds of capacity available while creates are failing. 3. **Make the failover path create nothing.** If it has to provision, promote or register, it depends on exactly the thing that is broken. 4. **Stop your own automation from shrinking you** while you cannot grow back. ## Reading the signal Tell the two apart from your own telemetry rather than from a status page, which usually lags the incident it describes. If your service's request success rate and latency are normal while your management calls error or hang, the control plane is degraded. Say that precisely: `the region is down` implies customers are failing and a whole footprint decision is on the table, while `the control plane is degraded` means customers are fine, your capacity is frozen, and the next decision is about what your own automation is allowed to do.

  • Why does the data path usually survive a control-plane incident at all?
    Because serving is executed by components that already hold everything they need - a running process, an attached disk, an installed route - and none of it requires asking a central service for permission to continue. The management API is the opposite: coordinated, region-wide and strongly consistent, so it has a much larger shared failure domain.
  • Name something that looks like ordinary runtime work but is really a control-plane call.
    Registering a freshly booted worker behind the entry point in front of a fleet; attaching a volume to an instance; a new instance reading its boot configuration from a management endpoint; changing a routing rule. Each is a change to platform state, so each fails in exactly the window where you most want it.
  • Why is the phrase 'the region is down' the wrong report here?
    It describes a different failure with a different response. A lost region means customer traffic is failing and you are deciding whether to move footprint. A degraded control plane means customers are still served, your fleet is frozen at its current size, and the decision is about what your own automation is allowed to do next.

Your badge still opens the office door and the lifts still run, but the building's leasing desk is shut: nobody can be issued a badge, given a room, or moved to a different floor.

saying these in an interview costs you the question

  • Says the region is down when only management calls fail
  • Thinks running instances stop when the console errors
  • Assumes a failover completes while create calls are hanging
  • Believes scaling will absorb the incident automatically
  • Treats the status page as the first reliable signal
open as a page

Your transcoding fleet's backlog is growing but every request to add a worker times out while running workers transcode normally - what now?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Accept a frozen fleet and manage demand instead of capacity: keep the durable queue absorbing arrivals, shed or defer the lowest-value work, freeze anything that could shrink or replace workers, and communicate delay rather than failure.

open as a page

How much pre-provisioned headroom should a platform team mandate for control-plane outages, and how do you justify paying for it?

level: principalimportance: should knowfreq 36%

basics

~20 s

Size headroom to the demand a workload must absorb while it can create nothing, not to average utilisation, and set it per tier rather than fleet-wide. Cheaper than any percentage: mandate that every failover path completes without creating a resource.

open as a page

During a management-API outage, which of your own automated loops can shrink a healthy transcoding fleet, and how do you stop them?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

Scale-in on a falling backlog, terminate-and-replace supervision, a deployment already in flight, and scheduled teardown can each destroy capacity you cannot rebuild. Suspend scale-in, switch replacement to drain-not-terminate, and pause deployments until creates succeed.

open as a page