skip to content

Why is Snowflake described as a hybrid of shared-disk and shared-nothing architecture?

level: middleimportance: must knowfreq 65%

answer

  1. two classical models, one of each half
  2. who owns the bytes decides the first half
  3. what happens inside one cluster decides the second
  4. immutability removes the old coherency problem
  5. shared storage plus MPP execution per warehouse

basics

~20 s

All compute reads one shared copy of the data in cloud object storage, which is shared-disk. Inside each virtual warehouse, nodes split the work with no shared memory or disk between them, which is shared-nothing MPP. Snowflake calls the result a multi-cluster, shared-data architecture.

solid answer

~50 s

The **shared-disk** half is the storage layer: every virtual warehouse in the account sees one authoritative copy of every table in cloud object storage, so there is no per-node ownership of data and no data movement when you add compute. The **shared-nothing** half is inside a warehouse: its nodes each take a disjoint set of files, process them against private CPU, memory and local SSD, and exchange data only through the query plan — classic MPP execution. Snowflake's own name for the combination is the *multi-cluster, shared-data* architecture. It matters because you get the elasticity of shared storage (spin up, resize or suspend compute with zero redistribution, run many warehouses on one dataset) together with the parallel scan throughput of shared-nothing execution. The price is that every cold read crosses the network to object storage instead of hitting an attached disk.

go deeper

for a junior

Recall that data sits in one shared place and that compute clusters are separate from it. You are not expected to compare classical architectures in depth yet.

for a middle

This is your tier: define shared-disk and shared-nothing, place each half of Snowflake correctly, and explain why immutable files in object storage sidestep the classic shared-disk bottleneck.

for a senior

Turn the model into operational reasoning — cold versus warm cache after suspends, why adding a warehouse never redistributes data, and where centralized coordination of concurrent writes shows up in practice.

for a principal

Frame the tradeoff at platform level: shared storage buys elasticity and single-copy governance, while network reads set a latency floor that no amount of tuning removes. Know which workloads that floor disqualifies.

## The two classical models **Shared-disk** systems put all data on storage every node can reach, and every node can serve any query. Adding a node adds compute immediately, with no data reshuffling. The historical weakness is the storage subsystem: every node hammers the same device, and keeping caches coherent across nodes requires a distributed lock or messaging layer that becomes the bottleneck. **Shared-nothing** systems give each node private storage and a slice of every table. Scans parallelize beautifully because each node reads its own local disk, and there is no cache-coherency traffic. The weakness is rigidity: the data placement is baked into the cluster, so scaling out means redistributing terabytes, and compute cannot be scaled or paused independently of the data. ## Where Snowflake sits Snowflake takes the shared-disk model for the *storage* layer and the shared-nothing model for *execution inside a warehouse*, and removes the classic weakness of each. **Shared storage.** All table data lives once in the cloud provider's object storage as immutable, compressed columnar files. Any virtual warehouse in the account can read any file. Object storage is effectively infinitely parallel and elastic — it is not the single spindle that made shared-disk fragile in the 1990s. And because the files are **immutable**, there is no cache-coherency problem to solve: a warehouse that has cached a file can never be holding a stale version of it. A write produces new files and a new table version in metadata; it never mutates a file another node has cached. **Shared-nothing execution.** When a warehouse runs a query, cloud services hands it a pruned list of files. The warehouse's nodes divide those files among themselves; each node scans, filters and partially aggregates against its own CPU, memory and local SSD cache, with no shared buffer pool. Data crosses node boundaries only where the plan demands it — joins and aggregations that require redistribution, and spilling when memory runs out. That is textbook MPP. ## What the hybrid buys - **Elastic compute with no data movement.** Resizing, suspending or creating a warehouse touches no table data, because compute owns none. In a pure shared-nothing warehouse, the same operations mean a redistribution. - **Many independent clusters over one dataset.** ELT, BI and ad-hoc analysis can run on separate warehouses that cannot contend for each other's CPU or memory, yet read the same rows with no copies, extracts, or replication lag. - **Scale-out scans.** Within one warehouse, throughput still rises roughly with node count for scan-heavy work, exactly as a shared-nothing engine promises. - **Independent scaling axes.** Storage grows without compute; compute grows without a data migration. ## What it costs - **Network on cold reads.** The first scan of a file goes over the network to object storage, with higher latency than a locally attached disk. The warehouse's local SSD cache hides this on repeat access — and is lost when the warehouse suspends. - **No data locality guarantees across queries.** Which node caches which file is not something you control, so warm-cache benefits are best-effort. - **Concurrent DML across warehouses is coordinated centrally.** Multiple warehouses writing the same table serialize through the transaction manager in cloud services rather than through per-node ownership. ## How to say it in an interview A crisp two-sentence version: "Storage is shared — one copy in object storage that every warehouse can read, so compute is stateless and elastic. Execution is shared-nothing — inside a warehouse, nodes split the file list and process their share privately, so scans scale with node count." Then add the reason the old shared-disk failure mode does not bite: immutable files plus cloud object storage remove both the coherency protocol and the single-device bottleneck. ## Common trap Candidates often say "Snowflake is shared-nothing like Redshift" or "Snowflake is shared-disk like an Oracle RAC cluster." Both halves are wrong on their own, and the interviewer is usually listening specifically for the hybrid framing and for *why* each half was chosen.

  • Why doesn't Snowflake need a cache-coherency protocol between warehouses?
    Because the stored files are immutable. A warehouse's local cache holds file chunks that can never change in place; a write creates new files and a new table version in metadata rather than editing an existing file. A stale cached file is therefore impossible — at worst a warehouse's cache holds files no longer referenced by the current table version.
  • What is the practical downside of reading from object storage instead of local disk?
    Latency and bandwidth on cold reads. The first scan after a resume, or of newly written data, crosses the network, so the same query can be visibly slower than its warm re-run. Frequent suspend cycles on a latency-sensitive workload trade credit savings for cold-cache penalties, which is why suspend policy is a tuning decision, not a default.

A commercial kitchen with one shared walk-in pantry: every crew can take from the same pantry (shared disk), but inside a crew the cooks split the tickets and work at their own stations with their own tools (shared nothing).

saying these in an interview costs you the question

  • Calls Snowflake plain shared-nothing like a classic MPP appliance
  • Says each warehouse node owns a fixed slice of the table
  • Claims a coherency protocol keeps warehouse caches in sync
  • Thinks resizing a warehouse redistributes stored data
  • Assumes shared storage means queries do not run in parallel

context