skip to content

Orchestration Model

What a cluster scheduler decides for you: where a workload lands, how many copies run, and what happens when one dies. Interviewers ask because running a container and running a fleet differ.

on this pageshow

questions

page 2 of 2

A change request moves an order service's workload spec from three copies to six and swaps its image digest — what can the reviewer conclude, and what must they check elsewhere?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The diff proves intent and nothing else: exactly these two fields are being changed, and the document is the complete statement of what is wanted. Whether six copies fit, what defaults apply unseen, and what is running today are all read elsewhere.

open as a page

On a shared cluster, when is dedicating a group of hosts to one workload class worth the capacity it strands?

level: principalimportance: should knowfreq 32%

basics

~20 s

Dedicated hosts are worth it when the workloads genuinely differ in the hardware they need, the interference they can tolerate, or the separation required of them — not when one team simply wants to be treated as more important.

open as a page

Six teams share one cluster and a seventh wants its own — how do you decide between sharing and one cluster per team?

level: principalimportance: should knowfreq 38%

basics

~20 s

Decide on three axes: the fixed floor each extra cluster carries, the blast radius teams are willing to share, and who absorbs each cluster's operational load. Separation inside one cluster is soft, so a security argument is a different question with a different answer.

open as a page

A batch job declares 500 required completions with 20 attempts running at once — what does the platform count, and when does it stop?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

It counts clean finishes, not started attempts. The at-once number is only a width limit: the platform keeps roughly that many attempts in flight, tops them up as each one ends, and marks the job finished when 500 have succeeded.

open as a page

What must a team's own controller do to reconcile a new declared object type the way the platform's built-in loops do?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A custom controller must register a declared type, be triggered by an object's identity rather than an event payload, re-read declared and live state on every pass, act idempotently, and report progress instead of failing terminally.

open as a page

On a managed control plane, which failures become the provider's problem and which stay yours?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

The provider runs the deciding half — the API, its state store and that store's backups — and is the only party who can repair it. Hosts, node agents, workloads and the declared state you put in stay yours.

open as a page

A batch group's work container exits successfully but its metrics helper keeps running — why does the group never finish?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Completion is a property of the group, not of one member. A run-to-completion group is done only once its members have stopped, so a helper with no exit condition keeps the group alive and the run never records success.

open as a page

A network fault makes twenty hosts stop reporting at once - what is the risk if the platform replaces everything they were running?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Mass replacement can turn an observation problem into a real outage: the surviving hosts cannot hold the whole fleet, every replacement pays a cold start at once, and if the hosts were never down, the cluster has just doubled its running workloads.

open as a page

showing 31–38 of 38