In a cloud warehouse, why is the first query after a suspended compute cluster resumes slower than later runs?
answer
- nothing on the new machines is warm yet
- resume means brand-new VMs
- the local SSD cache starts empty
- first run reads everything over the network
- the profile shows remote versus cached bytes
basics
~20 sResuming provisions fresh worker VMs whose memory and local SSD caches are empty, so the first query reads everything remotely from object storage. Subsequent runs hit warm local caches and skip most of that network traffic.
solid answer
~50 sTwo costs stack up. First, resume itself: the platform has to allocate and boot worker VMs and attach them to the cluster before any work starts. Second, and usually larger, every cache the cluster had is gone. A running cluster keeps recently read column chunks on local NVMe and in memory, so a repeat scan is mostly local; a freshly resumed cluster has nothing cached and must fetch every byte from object storage over the network, where per-request latency is far higher than a local disk read. Compiled-plan and metadata caches are cold too. The result is that a query which takes seconds warm can take many times that on the first run. Mitigations are to keep frequently used clusters alive across the working day, tune the idle timeout against the cold-start penalty rather than to the shortest possible value, or run a small warm-up query before a known schedule.
code
text · 9 lines-- first run after resume (cold)
scanned from remote object storage: 182 GB (100%)
scanned from local cluster cache: 0 GB (0%)
elapsed: 96s
-- identical query, immediately after (warm)
scanned from remote object storage: 6 GB (3%)
scanned from local cluster cache: 176 GB (97%)
elapsed: 11sgo deeper
Recall that a resumed cluster starts with empty caches and must fetch data over the network, so the very first query is the slow one.
Explain the components of the penalty — VM provisioning, empty local SSD and memory caches, cold metadata and plan caches — and say which one usually dominates.
Diagnose it from a profile: show that you check the remote-versus-cached byte split before blaming plans or statistics, and that you tune idle timeouts against arrival patterns.
Frame it as the price of statelessness and set policy: which workloads get kept warm, which get pre-warmed on a schedule, and how you justify idle compute spend against latency targets.
## Why a cold cluster is slow When storage and compute are separated, a compute cluster is stateless with respect to persistent data — it holds caches and nothing else. That is what makes suspend and resume cheap for the platform, and it is exactly what makes the first query after resume expensive for you. Three things are cold at the moment a cluster comes back. **The machines.** Resuming means allocating worker VMs, starting the engine process on each, and joining them into a cluster. This is real wall-clock time before a single byte is read. On most platforms it is seconds to tens of seconds depending on cluster size and whether the provider keeps a warm pool; treat the exact figure as something you measure for your own platform rather than something you quote. **The local data cache.** A running cluster stores recently read column chunks on local NVMe SSD and in page cache. On a repeat scan most reads are satisfied locally at microsecond-to-low-millisecond latency. A resumed cluster's SSDs are empty — the old ones belonged to VMs that no longer exist — so every column chunk must come from the object store over the network, where per-request latency is orders of magnitude higher than local NVMe and throughput is bought by issuing many parallel ranged reads. The scan is not slower because the CPU is slower; it is slower because it is I/O-bound on a remote store. **Everything else warm-path.** Metadata for the tables being queried has to be fetched and parsed again, and engines that cache compiled plans or query fragments lose those too. These are smaller terms than the data cache but they show up on the very first statement. ## Reading it in the profile Most engines report, per query, how much data was read from remote storage versus served from the local cache. That split is the diagnostic. If the first run shows near 100% remote and the second run shows the same scan mostly local, you are looking at cache warm-up, not a plan regression, not a statistics problem, and not a data-layout problem. If instead the second run is also mostly remote, the working set is larger than the cluster's local cache and no amount of warming will fix it — you need a bigger cluster (more aggregate local SSD), better pruning so the working set shrinks, or acceptance that this workload is remote-read-bound. ## What resizing does Resizing is a partial version of the same effect. Adding workers gives you machines with empty caches and, on engines that assign files to nodes consistently, changes which node is responsible for caching which file — so even the surviving nodes' caches become partly useless. Removing workers destroys their cached data outright. A cluster that has just been resized will behave cold for a while even though it never suspended. ## Managing the trade-off The idle timeout is the control. Suspend aggressively and you stop paying for idle compute but pay cold-start on the next query; suspend lazily and queries stay fast while you burn compute on an empty cluster. The right setting follows the arrival pattern, not a rule of thumb: - **Interactive BI during business hours** — arrivals are bursty but continuous, so a short timeout guarantees users hit cold starts all day. Keep the cluster alive through the working window; the idle cost is small relative to the number of queries protected. - **A single nightly batch job** — nothing is served interactively, one cold start amortizes over an hour of work, so suspend as fast as the platform allows. - **A scheduled morning dashboard refresh** — schedule a cheap warm-up query a few minutes ahead so the cluster is up and the hot tables are partly cached when real users arrive. Two common mistakes are worth naming. Shortening the idle timeout to save money on a bursty interactive workload usually costs more than it saves, because every resume re-reads the working set remotely and the queries themselves run longer, burning compute-seconds. And downsizing the cluster does not shorten cold start meaningfully — a smaller cluster boots marginally faster but has less aggregate local SSD, so it stays cold longer and holds less of the working set. ## The framing that lands in an interview Cold start is the bill for statelessness. The same property that lets you suspend a cluster to zero cost, resize it without moving data, and run several clusters over one copy of the data is the property that guarantees a resumed cluster knows nothing. You do not eliminate it; you decide where to spend — idle compute, or first-query latency.
- How would you set the idle-suspend timeout for a cluster that serves interactive dashboards through the business day?Not to the minimum. Bursty interactive arrivals mean a short timeout produces repeated cold starts, and each one re-reads the working set remotely and lengthens every query. Keep the cluster alive across the working window and suspend outside it, then compare the idle compute cost against the query-time cost you avoid.
- Does a cluster's local cache survive a resize?Only partly, and sometimes not at all. Removed workers take their cached data with them, added workers arrive empty, and if the engine assigns files to nodes consistently, the reassignment invalidates much of what the surviving nodes cached. Expect cold-ish behaviour after any resize, not just after a resume.
- The second run of a query is still mostly remote reads. What does that tell you?The working set exceeds the cluster's aggregate local cache, so entries are evicted before they are reused. Warming will not help. Shrink the working set with better filtering and pruning, run a larger cluster for more aggregate SSD, or accept that the workload is bound by remote-read throughput.
saying these in an interview costs you the question
- Blames a plan regression rather than an empty cache
- Thinks a smaller cluster resumes fast enough to fix it
- Claims the local cache survives suspension of the cluster
- Sets the idle timeout to the minimum for interactive workloads
- Confuses the result cache with the cluster's local data cache