Your platform runs an always-on checkout service and a nightly report that must finish once — why are those different workload contracts?
answer
- what does a clean exit mean here?
- live copies against successful runs
- an exiting service is a hole to fill
- a finished job is a terminal state
- four shapes: always-on, finish, per-host, recurring
basics
~20 sThey disagree about what a process exiting means. A replicated long-running service must never exit, so the platform recreates any copy that does. A run-to-completion job is supposed to exit once; a clean exit is the whole point and ends it.
solid answer
~50 sA cluster platform does not merely start containers — it holds an invariant, and the workload shape you declare is which invariant you asked for. For a replicated long-running service the invariant is `this many copies are running right now`, so a copy that exits, cleanly or not, is a hole the platform fills. For a run-to-completion job the invariant is `this many successful runs have happened`, so a clean exit increments a count and, once the count is reached, nothing is started again. Two more shapes sit alongside them: a one-copy-per-host workload, whose count is derived from the set of matching hosts rather than declared, and a recurring schedule, which is a factory that starts a run at each trigger time. Picking the wrong one does not fail loudly — it produces work done over and over, or work that quietly stops.
code
yaml · 11 lines# the same program, declared under two different contracts
- name: checkout-api
contract: long-running-service
copies: 4 # the platform holds this many, always
# a copy that exits is replaced, clean or not
- name: nightly-reconcile
contract: run-to-completion
requiredCompletions: 1 # one clean finish and the workload is done
retryBudget: 3 # a failed attempt may be started againgo deeper
Recall the fork: a replicated service must never exit and any copy that does is replaced, while a run-to-completion job is supposed to exit once and a clean exit ends it. Say that out loud before naming anything else.
Explain the invariant behind each shape — a count of live copies against a count of successful runs — and then derive the consequences: how the size is changed, what an update means, and why zero running is an emergency for one and normal for the other.
Show the operational half. Name what a job's finished record buys you at 8am, what alerting each shape deserves, and describe a workload you re-declared under a different contract after it misbehaved in production.
Frame it as a platform policy question: which contract is the default for a new workload, who decides, and what guardrail keeps a batch from being shipped as an always-on service. The cost of a wrong default is paid nightly across every team.
## Four contracts, one platform Running a container is a verb; declaring a workload is a **contract**. You hand the platform a document describing an end state, and the platform keeps acting until reality matches it. What differs between workload shapes is *which* end state is being held, and that single choice decides what an exiting process means, how you change the size of the thing, and what your alerting should page on. | Shape | The invariant the platform holds | Does a copy ever end normally? | |---|---|---| | Replicated long-running service | this many copies exist **right now** | no — an exit is a fault to repair | | Run-to-completion job | this many **successful runs** have happened | yes — a clean exit is the goal | | One copy per host | exactly one copy on **every matching host** | no — the count follows the host set | | Recurring schedule | a run was **started at each trigger time** | each run ends; the schedule does not | ## What happens the moment a process exits This is the fork the whole question turns on. - Under the **service** contract, the platform compares running copies against the declared number, sees a shortfall, and creates another copy. It does this whether the process ended in an error or returned quietly, because the contract says nothing about *why* — only that the number is wrong. - Under the **job** contract, the platform asks a different question: did that run finish its work? A clean finish increments the count of completions. When the declared number of completions is reached, the workload is **finished** — a terminal state with a recorded result. An unclean finish is a failed attempt, which the platform may start again until a retry budget is used up. - Under the **one-copy-per-host** contract, the count is never declared at all. It is derived: one copy per host that matches, so adding a host adds a copy and removing a host removes one. - Under a **recurring schedule**, the schedule itself runs forever and each run it creates is an ordinary run-to-completion job with its own outcome. ## What follows from the choice 1. **How you change the size.** A service has a copy count you set. A job has a required number of completions and how many may run at once. A one-copy-per-host workload has neither — you change which hosts match. 2. **What an update means.** A service is replaced stepwise while it keeps serving. A job is not updated in flight; you change the declaration and run it again. 3. **What alerting should say.** `Zero copies running` is an emergency for a service and completely normal for a job between runs. Teams that alert on the same signal for both either page all night or never page at all. 4. **What history exists.** A job leaves a record — started, finished, succeeded or failed, how long it took — that a human or another system can read afterwards. A service's finished copies are simply gone; what you have instead is its current state. ## The nightly batch, concretely A reconciliation batch has to finish before morning and shares the platform with always-on services. Declared as a job, the platform knows what "done" means, keeps the result, retries an attempt that failed, and stops when the work is done. Declared as a service, that same batch would be started again every time it succeeded: the platform sees a copy that stopped existing, and dutifully recreates it, running the reconciliation over and over all night. The mirror mistake is quieter. A long-running consumer declared under the job contract exits cleanly the first time it has nothing to do — and the platform records success and moves on. No copy is missing, because the contract never promised one. ## What an interviewer is listening for Not the names of the shapes. They want the sentence "what does a clean exit mean here?", because that single question separates the four contracts and explains every symptom that follows from getting it wrong. A candidate who answers with the invariant — *the platform is holding a count of live copies* versus *a count of successful runs* — has the model. A candidate who answers with a list of object names on one platform has memorised a menu.
- Where does a recurring schedule fit among the four shapes?It is not a fifth kind of running thing — it is a factory. At each trigger time the schedule creates a run-to-completion job and that job carries the outcome. The schedule holds one invariant of its own: that a run was started for each trigger time it was supposed to fire, subject to whatever it does when a run is late or when the previous one is still going.
- A batch job succeeds every night, so why keep a record once it is finished?Because the finished state is the only evidence the work happened. The record carries when the run started, how long it took, whether it succeeded and how many attempts it needed — which is what you read when this morning's numbers look wrong. A service has no equivalent: its old copies leave nothing behind, so what you keep for a service is its live state and its shipped output instead.
A thermostat holds the room at a temperature and never stops; an oven timer is supposed to ring once and then be done. Asking a thermostat to ring once, or a timer to hold a temperature, is the same category error.
saying these in an interview costs you the question
- Says a job is just a service that happens to stop on its own
- Thinks the platform decides the shape from how long the process runs
- Claims a service that exits cleanly is left alone by the platform
- Assumes a copy count means the same thing for every workload shape
- Alerts on zero running copies for a batch workload and calls it monitoring
- Says you pick the shape for convenience and nothing downstream depends on it