When a job's input lives in remote blob storage, what happens to the scheduler's preference for nearby bytes?
answer
- nothing to prefer when all are equal
- traded away deliberately, not lost
- read less rather than place better
- region boundary is the surviving unit
- affinity survives inside the run only
basics
~20 sIt disappears, on purpose. No compute machine holds the bytes, so every candidate is equally far and placement collapses to free capacity. The industry traded that optimisation for compute it can resize or discard independently of the data.
solid answer
~50 sRemote blob storage is reached over the network and runs on machines a piece of work cannot be placed on, so there is no nearer candidate to prefer: every lane in the cluster is equidistant from every byte. The scheduler's nearness preference therefore has nothing to express, and the sensible policy becomes first free lane. That was not an accident — it is what buying elastic, disposable compute costs. What remains is coarse: keep the compute in the same region or availability zone as the store, because crossing that boundary changes both bandwidth and price, and expect aggregate read throughput to be bounded by the network path to the store rather than by disks. Some platforms add a cache on the compute side that re-creates a weak affinity; that layer belongs to the warehouse architecture subject, not here.
go deeper
Recall that when data lives in remote storage, no machine in the cluster holds it, so there is no nearer machine to pick. Placement becomes simply whichever lane is free.
Explain the consequence: every input byte crosses the network, so reading fewer bytes replaces placing work well as the thing worth optimising.
Argue both sides of the trade. Name what independence bought, name the bandwidth ceiling it imposed, and identify the region boundary as the only input placement decision that survived.
Treat it as a platform posture. Decide whether any workload you own still justifies storage on the compute machines, and stop paying for nearness machinery on the ones that do not.
## What the architecture changed An **object store** is remote blob storage reached over the network: it runs on machines that are not part of your cluster, and there is no machine in your cluster that a piece of work could be placed on to make a read local. That single fact removes the input side of placement entirely. Recall what the preference needed in order to mean anything. It needed a **cluster file system** — storage running on the same machines as the compute, with each block on the local disks of a few of them — so that some candidate machines were genuinely nearer to a given slice than others. Take the storage off those machines and the ordering collapses: every lane is the same distance from every byte, so there is no rung to fall through and no wait that could pay for itself. Placement becomes a capacity question: which lane is free. ## This was bought, not lost by accident The trade was made deliberately, and the thing bought was independence: - Compute can be resized, suspended or destroyed without touching the data, because the data was never on it. - Storage grows without buying processors, and processors are bought without growing storage. - A machine's death costs no data at all — there is nothing stored on it that anyone else needs. - Several clusters can read the same bytes at once, since no cluster owns them. The detailed economics and machinery of that architecture — elastic compute, the caching layers bolted onto it, table metadata and transactions over passive storage — belong to the warehouse-architecture subject. What belongs *here* is the one consequence: **it deleted nearness over input bytes as a scheduling concept.** ## What it costs on the read path | before: co-located storage | after: remote storage | |---|---| | some machines nearer, most reads local | all machines equidistant, every read remote | | read bandwidth scales with local disks added | read bandwidth bounded by the network path to the store | | scheduler models rungs and waits | scheduler places on the first free lane | | losing a machine can lose the only nearby copy | losing a machine loses no data | The practical consequences on the second column: 1. **Every input byte crosses the network.** Reading less therefore matters much more than placing well — filtering before the read, and reading fewer files, are what replaced placement as the lever. 2. **Request behaviour becomes a first-class cost.** Many tiny objects read one at a time perform quite differently from the same bytes in fewer, larger objects, and this is a property of the store and its protocol rather than of any engine. 3. **Aggregate bandwidth is shared.** Adding lanes raises demand on a path that is not getting wider, so past a point more parallelism buys nothing. ## The nearness that survives, at a coarser grain Placement did not vanish so much as change units. What is left is a boundary question, not a machine question: - **Same region, and where it applies, same availability zone as the store.** Crossing that boundary changes latency, achievable bandwidth and, on most platforms, the price per byte moved. This is the single placement decision still worth making about input. - **Inside a running job**, affinity does survive in two specific places — bytes a machine produced for itself, and work bound to a key whose accumulated state one worker already holds. That is a different subject from input placement, and it is the one place where "which machine" still has an answer. ## Where deployments genuinely differ - Not every deployment moved. Clusters with storage on the compute machines still exist where read latency, data movement cost or regulatory locality dominate, and there the original preference still pays. - Some platforms keep a compute-side cache of remote bytes, which re-creates a weak preference for the machine that read an object before. Whether an engine exploits that, and whether the cache survives a cluster being resized, varies; the cache layer itself is owned by the warehouse-architecture subject. - Runtimes differ in whether they still express a nearness preference at all in this environment. Where one does and no machine is nearer, the configured wait is a pure cost — idle lanes bought for a saving that cannot exist. ## Answering it in an interview The weak answer is "data locality is a best practice". The strong answer is that it was a best practice for one architecture, that the industry knowingly traded it for compute it could throw away, and that the levers which replaced it are reading fewer bytes and staying inside the store's own region — with real, surviving affinity found only inside a run, not on its input.
- Does any placement decision about input still matter in that architecture?One: keep the compute in the same region, and where the store exposes them the same availability zone, as the data. That boundary changes latency, achievable bandwidth and usually the price of the bytes moved. Below that boundary, no machine is nearer than another, so nothing finer is worth expressing.
- Why was this trade made deliberately rather than regretted?Because compute that holds no data can be resized, paused or destroyed freely, and storage can grow without buying processors. The nearness optimisation was worth a scheduler's complexity only while networks lagged local disks badly; once that gap narrowed, independence was worth far more than the remaining read saving.
- If a cluster keeps a local cache of bytes it has read, has nearness returned?Only weakly and only after the first read. It creates a preference for the machine that read an object before, which lasts as long as that machine and its cache do. It is an optimisation on top of remote storage rather than a return to co-located storage, and the cache layer belongs to the warehouse-architecture subject.
saying these in an interview costs you the question
- Insists local placement still helps when input is in remote storage
- Calls the separation a regression rather than a deliberate trade
- Thinks the store reports which compute machine holds each object
- Believes adding lanes always increases total read throughput
- Treats region placement and machine placement as the same decision