Eight teams submit a hundred jobs a day, from seconds long to hours long. How would you place those workloads across the three supply models?
answer
- per workload, not per organisation
- wait tolerance versus lifecycle ownership
- who is on call for the machines
- each extra model costs a runbook
- say what your platform actually offers
basics
~20 sPlace by wait tolerance and lifecycle ownership, not by team. Short interactive work needs machines already up; long scheduled jobs earn their own cluster and their own runtime versions; work with no operator behind it fits a service that shows no machine. Then limit how many models you operate.
solid answer
~50 sDo not pick one model for the organisation; pick per workload, on three axes. First, **how long the work can wait before it starts** — raising machines for a job that runs for seconds is mostly waiting, so short interactive work belongs somewhere already warm. Second, **whose lifecycle it needs** — a long, heavy, scheduled job earns machines of its own, because it gets its own runtime versions, its own blast radius and no leftovers. Third, **who is available to operate it** — a team with no on-call and unpredictable arrival is better served by a service that shows no machine and owns sizing itself. Then apply the counterweight: every model you support adds an access path, a sizing conversation, a failure mode and a runbook, so the number of models in use is itself the decision you are making.
go deeper
Know that the choice exists and is made per workload: some work waits badly and needs machines that are already there, and some work is long enough that the wait does not matter.
Argue the trade in both directions — the start-up wait against the isolation and independent versions a cluster of your own gives you — and say which workload shapes land on each side.
Bring the operational axes in: who is on call for the machines, how predictable arrival is, whether anything needs to stay warm, and what a never-ending job does to a shared cluster.
Make it a policy question. Pick the smallest set of models that covers the real spread, name the default, and count the runbooks, access paths and failure modes each additional model adds.
## The question is per workload, not per organisation The weak answer picks a favourite and standardises on it. The strong answer observes that a hundred jobs a day across eight teams is not one workload, and that the three supply models differ on exactly the axes those jobs differ on. Two definitions the argument needs: a **worker process** is the process on a cluster machine that runs pieces of a job and owns their memory, and a **coordinating process** is the single process that plans the pieces, hands them out and tracks what finished. Every model produces both; the models differ in who owns the machines under them and for how long. ## The axes that decide a placement 1. **Tolerance for the wait before the first piece runs.** Machines that already exist start work now. Machines that must be raised first spend a fixed wait doing so, and how long that is varies widely by platform. For a job that runs for hours, the wait is a rounding error. For work someone is sitting and waiting on, it can be most of the elapsed time. 2. **Whose lifecycle the workload needs.** A job that wants its own runtime and library versions, its own restart schedule and a blast radius of exactly itself wants its own machines. A job that is happy with whatever everyone else has does not need to pay for that. 3. **Who will operate it.** Keeping machines up is an ongoing job with an owner. A team without one is better off with a model where the provider owns sizing, starting and stopping. 4. **How predictable the arrival is.** Steady, known demand is what a long-lived pool is sized against. Spiky, unpredictable demand from many teams is what it is worst at, because it is sized in advance for everyone at once. 5. **What the workload keeps.** If the second run genuinely benefits from a warm working set, machines that die with the job cannot provide it; you must either keep machines up or persist the result to shared storage deliberately. ## A placement that usually falls out | workload shape | model | why | |---|---|---| | short interactive or exploratory work, all day | machines already up | the wait to raise machines would dominate the work itself | | long scheduled jobs, heavy, predictable | a cluster per job | isolation and independent versions are worth the start-up wait | | continuous jobs that never end | a cluster per job | one owner, one lifecycle, no neighbours for something that runs forever | | occasional work from a team with no operator | a service that shows no machine | nobody has to own sizing, patching or restarts | | anything odd, rare and one-off | a service that shows no machine | the least standing commitment for the least frequent work | The continuous-job row is worth defending in an interview: a job that never ends does not mix well with machines shared by everyone, because it holds its share forever and it is the job most disrupted by a restart of the shared thing it lives on. ## The counterweight nobody volunteers Each model you operate has a cost that is not machine time: - a separate way of submitting work, and a separate thing that goes wrong on submission; - a separate access and permission path; - a separate place logs and run history end up; - a separate sizing conversation to have with each team; - a separate runbook, and a separate set of failures an on-call person must recognise at three in the morning. So the principal's move is not to place every workload optimally; it is to pick **the smallest set of models that covers the real spread**, and to make the default obvious so teams do not have to choose. Two is usually the honest answer for an organisation of this size — something warm for interactive work, and per-job machines for everything scheduled — with the third reserved for the long tail if it is already available. ## What varies, and must be said out loud The platform decides much of this for you, and the market genuinely differs. On some platforms a cluster is raised in well under a minute and the wait argument nearly disappears. On others a job is sized from a hint rather than an exact request, so "we will size it ourselves" is not on offer. Some services accept no machine sizing at all, only a coarse capacity tier. Any recommendation that does not say which of these is true of your platform is a recommendation about somebody else's platform. ## What this decision does not settle It does not settle how a finite pool is divided when several jobs want it at once, nor what a pool costs while nobody is submitting to it. Those are separate calls with separate owners, and conflating them with the supply model is how placement discussions turn into budget arguments.
- Why do continuous jobs that never end sit awkwardly on machines shared with everyone else?Because they hold their share indefinitely rather than finishing and giving it back, and because a restart or upgrade of the shared machines interrupts something that was supposed to run forever. Their own machines give them one lifecycle and one owner, which is what a never-ending job needs.
- A team argues for standardising every workload on one model. What is the strongest counter?That the workloads differ on the axes the models differ on. Short interactive work cannot absorb the wait to raise machines, and a long isolated job does not benefit from sharing. Standardising is still worth something, so the real answer is the smallest set of models that covers the spread, not one.
- What changes your recommendation if raising a fresh cluster on your platform takes only a few seconds?Most of the case for keeping machines up disappears, since the wait argument was the main one. What remains is reuse of warm data across jobs, which per-job machines cannot provide. That is a narrower reason, and often better solved by persisting the result to shared storage.
saying these in an interview costs you the question
- Picks one supply model for the whole organisation with no per-workload argument
- Puts short interactive work on machines that must be raised first
- Ignores who will operate the machines once they exist
- Treats the number of models operated as free
- Recommends a model without saying what the platform actually offers