On a managed control plane, which failures become the provider's problem and which stay yours?
answer
- the line sits at the API's edge
- they operate it, you cannot repair it
- the record's contents are still yours
- hosts and agents usually stay yours
- your only lever is preparation
basics
~20 sThe provider runs the deciding half — the API, its state store and that store's backups — and is the only party who can repair it. Hosts, node agents, workloads and the declared state you put in stay yours.
solid answer
~40 sThe line is drawn at the API's edge. Everything behind it — the API itself, the state store, its backups, and version upgrades of those components — is operated by the provider, which is the whole appeal: there is no store to size, patch or restore. Everything in front of it stays yours: the hosts and node agents in most offerings, the workloads, their data, and the *contents* of the declared state, because a bad spec you applied is still a bad spec. The real trade is leverage. When the deciding half is unavailable you cannot restart it, roll it back or fail it over; your only lever is the preparation you did beforehand — enough copies already placed and spread, and workloads that do not call the cluster API while serving a request.
go deeper
The basic idea: someone else runs the part that decides, and you interact with it only through its API. What runs your workloads is still yours to operate.
Be able to sort concerns onto the right side of the line — the store and its backups theirs, the declared state's contents, the hosts and the workloads yours.
Show what changes in an incident: you cannot restart, roll back or restore their half, so everything you can do about it was decided in advance by how the workloads were placed and written.
Frame it as buying operational surface at the price of leverage, and set the estate-wide standards that make the trade safe: exportable declared state, spread capacity, and no cluster-API dependency on any request path.
## Where the line is drawn A managed control plane means someone else operates the deciding half and exposes only its API to you. The split is clean to describe and easy to get wrong in an incident. | Concern | Managed offering | Self-operated | |---|---|---| | The API's availability and capacity | Provider | You | | The state store: sizing, growth, backups, restore | Provider | You | | Version upgrades of the deciding half | Provider, usually on their schedule | You | | Hosts and node agents | You, in most offerings | You | | The declared state you applied | You | You | | Workloads, their data and their dependencies | You | You | | Identity and permission rules at the API | You configure, provider enforces | You | ## What genuinely moves - **Operating the store.** No growth to watch, no restore to rehearse, no capacity plan for the record itself. - **Patching and upgrading the deciding half.** It happens without you building the runbook. - **The deciding half's own resilience.** Spreading its components and recovering them is the provider's engineering problem, not yours. ## What never moves - **The contents of the record.** A spec that declares the wrong image, the wrong count or the wrong permissions is yours, and no managed offering examines it for you. - **Your hosts and their agents**, in most offerings — including whether their version is compatible with the deciding half the provider just upgraded. - **The workloads and their data.** A managed deciding half restores nothing about your volumes, your databases or your images. - **Designing for the outage hour.** Which is the real point. ## The leverage you trade away Self-operating means that when the deciding half is unavailable you can act: restart a component, roll a version back, restore the store, add capacity. A managed control plane removes both the work and the option. During an outage of the provider's half your team has exactly one lever, and it was pulled weeks earlier: 1. **Enough copies, already placed and already spread**, so losing one host during the window costs latency rather than availability. 2. **No cluster-API call on the request path**, so an outage of the deciding half cannot become a user-visible one. 3. **Capacity provisioned ahead of a known peak**, because scaling is the first thing that stops. That is not an argument against managed offerings — for most teams the provider operates the deciding half better than they would. It is an argument for being clear-eyed that *you did not remove the outage hour, you outsourced it*, and the time you got back should partly be spent on the three items above. ## Questions worth asking before choosing - **Can you export the declared state yourself**, so your intent exists outside the provider's store? - **What is the restore path**, and what does the provider restore — the record only, or anything about workloads? - **How much version skew is supported** between the deciding half and the node agents and clients you run, and who schedules the upgrade? - **Which failures are visible to you at all?** A degraded deciding half that is slow rather than down is the hardest case, and you cannot inspect it. ## Where designs differ Offerings vary widely in how much of the serving half they also take. Some leave you every host and agent; some run the hosts too and hand you only a place to put workloads. The further that line moves, the less there is to operate and the fewer levers remain — the same trade, applied to the other half. Name the trade rather than a product, and the answer holds whichever offering the interviewer has in mind.
- What is your only real lever during an outage of a managed deciding half?Preparation, all of it done earlier: enough copies already placed and spread across hosts, capacity provisioned ahead of the peak, and no cluster-API call on the request path. During the hour itself you can watch and communicate, but there is nothing in the deciding half for you to fix.
- When the provider upgrades the deciding half, what does that put at risk on your side?The version gap between the deciding half and the node agents and client tools you still operate. Platforms publish a supported skew for exactly this reason, and the usual break is a spec field your agents do not understand yet, or one that was removed underneath a tool you depend on.
saying these in an interview costs you the question
- Says a managed control plane removes the outage hour entirely.
- Thinks the provider is responsible for a bad spec you applied.
- Assumes every managed offering also operates the hosts and node agents.
- Believes managed backups remove the need to keep declared state elsewhere.
- Treats a managed deciding half as a reason to stop spreading copies.