A job ran on machines raised only for it and cached a working set on them. What survives after those machines are torn down?
answer
- the disks go with the machines
- what lived elsewhere survives
- logs must leave before teardown
- every run starts cold
- recovery points belong off-cluster
basics
~20 sOnly what was written to storage outside those machines. Caches held in worker-process memory, intermediate output written to local disk, temporary directories and logs left on the machines all disappear with them, so the next run starts cold.
solid answer
~50 sNothing that lived on the machines. When a cluster is raised for one job and torn down at the end of it, the machines and their local disks go together, so anything held in a worker process's memory — the process that runs pieces of the job and owns their memory — and anything written to local disk is gone. What survives is exactly what the job wrote somewhere else: results in shared storage, and logs or run history only if something shipped them off before teardown. The practical consequence is that a repeated job on this model never reuses the previous run's warm data: every run re-reads its input and rebuilds any cached working set. A long-lived shared pool can keep some of that across jobs, which is one of the few things it actually buys.
go deeper
Remember the rule of thumb: if it lived on those machines, it is gone. Only what the job wrote to storage somewhere else is still there afterwards.
Explain the mechanics — memory-held working sets, data written to local disk when memory ran short, temporary directories and logs all share the machine's lifetime — and say what that costs a job that runs every night.
Show the operational habits that follow: stream logs off the machines during the run, persist intermediate results deliberately, and keep recovery points in storage that outlives the cluster.
Weigh the reuse a long-lived pool makes possible against the leftovers it accumulates, and decide which of the two your workloads actually care about before standardising on a model.
## What teardown removes When machines are raised for one job and released when it finishes, the release is not selective. It takes the machines, the processes on them and the disks attached to them. Concretely, all of this goes: - **Anything in a worker process's memory.** A worker process is the process on a cluster machine that runs pieces of the job and owns the memory those pieces use. A working set deliberately kept in that memory for reuse is kept for the life of that process, and that process has no life after the job. - **Anything written to a machine's local disk.** That includes data a worker wrote out because it did not fit in memory — spilling, meaning writing part of a working set to local disk to keep going with less memory — and, on engines whose design materialises intermediate output before it is fetched, that intermediate output too. - **Temporary directories, unpacked program archives, downloaded dependencies and anything a start-up script left behind.** - **Logs and any run history held only on the machines**, which is the one that actually hurts, because you usually want them precisely when the job failed. ## What survives, and only because it lived elsewhere The surviving set is easy to state: whatever the job put somewhere that is not those machines. | artefact | survives teardown? | why | |---|---|---| | final output in shared or remote storage | yes | the storage has its own lifetime | | a working set cached in worker memory | no | the process holding it is gone | | data written to local disk when memory ran short | no | the disk is gone with the machine | | logs streamed to a collector during the run | yes | they left the machine before teardown | | logs left in a directory on the machine | no | nothing copied them off | | a saved recovery point written to shared storage | yes | it was written off-cluster on purpose | That last row matters for a job that is meant to be resumable. A saved recovery point — a consistent picture of the job's progress, written so a restart can continue from it rather than from the beginning — is only useful if it was written somewhere with a longer life than the machines. On this supply model that is not a nice-to-have; it is the difference between a restartable job and one that always begins again. ## Where engines genuinely differ Do not over-specify what is on the local disks in the first place, because the family disagrees. - Engines in the older batch lineage deliberately write intermediate output to local disk and have consumers fetch it afterwards, so a torn-down machine can take a lot of completed work with it. - Engines that push records across the network as they are produced materialise far less locally, so there is less on those disks to lose — but they lean harder on a saved recovery point, which had better be off-cluster. - How much is cached at all is a property of the program, not of the supply model: a job that never asked for a working set to be held loses nothing by being given fresh machines every time. So the honest statement is: what the machines held varies by engine and by program; that whatever they held is gone does not vary at all. ## What this changes about how you work 1. **Assume every run is cold.** If a pipeline's second step was fast last time because the first step's output was still warm on the machines, that saving does not exist here. Persist the intermediate result to shared storage and read it back, or accept the recomputation. 2. **Ship logs during the run, not after.** Anything you want to debug with must leave the machine while the machine exists. A job that failed and took its own logs down with it is a common and entirely avoidable outcome on this model. 3. **Write recovery points off-cluster.** Otherwise a restart cannot use them. 4. **Stop treating local disk as durable.** It is scratch space with the lifetime of a machine, and on this model the machine's lifetime is one job. ## The contrast that makes it a question A long-lived shared pool is the mirror image: the machines outlive the job, so a cached working set can in principle be reused by something submitted later, scratch directories accumulate, and logs sit where they were written until someone cleans up. That reuse is one of the genuine advantages of keeping machines up, and it comes with the matching disadvantage that one job's leftovers are still there for the next one to trip over. A managed compute service shows you no machine at all, so the question does not arise in the same form: there is nowhere for you to have left anything.
- The same job runs nightly on fresh machines each time. What can you do to stop it recomputing the same intermediate result?Write the intermediate result to shared storage as a real output and have later runs read it back. Caching it in worker-process memory cannot help, because the processes are new every night. The trade is an extra write and read against the recomputation you avoid.
- Why is a saved recovery point written to a machine's local disk close to useless on this supply model?Because the point of a recovery point is to survive the failure you are recovering from, and here the failure that ends the cluster also removes the disk holding it. It has to be written to storage with a longer lifetime than the machines for a restart to find it.
- Does a long-lived shared pool guarantee the next job can reuse a cached working set?No. Reuse is possible because the machines are still there, but the cache lives inside a worker process, and whether that process survives, still holds the data, and is even allocated to the next job depends on the platform. It is an opportunity, not a guarantee.
saying these in an interview costs you the question
- Expects a cached working set to be reused by the next run
- Assumes a finished job's logs can be read off machines that no longer exist
- Treats a machine's local disk as durable storage
- Writes recovery points to local disk on machines raised per job
- Thinks teardown removes only the job's processes, not their disks