skip to content

What are the three layers of Snowflake's architecture, and what does each one own?

level: juniorimportance: must knowfreq 85%

answer

  1. three concerns, three independent bills
  2. the bytes never live on the compute nodes
  3. compute clusters have a name of their own
  4. the brain above both is multi-tenant
  5. storage, virtual warehouses, cloud services

basics

~20 s

Snowflake separates storage (compressed columnar data in cloud object storage), compute (virtual warehouses, independent clusters that run queries), and cloud services (metadata, query optimization, transactions, security). The three layers scale independently and are billed separately.

solid answer

~40 s

Snowflake is a **three-layer** system. **Database storage**: loaded data is rewritten into Snowflake's own compressed columnar files and kept in cloud object storage (S3, Azure Blob, or GCS) in the account's region; you only reach it through SQL. **Query processing**: a *virtual warehouse* is a cluster of compute nodes with its own CPU, memory and local SSD cache that reads what it needs from the storage layer. Many warehouses can run against the same tables at once without contending for each other's resources. **Cloud services**: the multi-tenant control plane that owns authentication and role-based access control, the catalog and file-level metadata, query parsing and optimization, transaction management, and the result cache. Because storage and compute are decoupled, you can grow data without buying compute, and add or resize a warehouse without moving a byte.

code

sql · 7 lines
sql
-- Two independent compute clusters over the SAME stored table
USE WAREHOUSE etl_wh;
INSERT INTO sales SELECT * FROM staging_sales;

-- different session, different warehouse, no contention
USE WAREHOUSE bi_wh;
SELECT region, SUM(amount) FROM sales GROUP BY region;

go deeper

for a junior

Be able to name the three layers and one responsibility of each without hesitating. Knowing that data sits in cloud object storage and that compute clusters are called virtual warehouses is the bar here.

for a middle

Explain what each layer is billed for and why storage and compute scale separately. Expect to be asked which layer holds metadata, which holds the result cache, and what a warehouse actually caches locally.

for a senior

Show you use the split operationally: sizing and isolating warehouses per workload, reasoning about cold-cache behaviour after a suspend, and knowing that pruning decisions happen in cloud services before the warehouse reads anything.

for a principal

Own the consequences at platform scale — warehouse topology per team, credit accountability per layer, and where the shared multi-tenant services layer becomes a coupling point across an organization's accounts.

## Why the split exists Classic warehouse appliances bind data to the machine that stores it: each node owns a slice of every table, so adding compute means redistributing data, and one heavy workload starves every other one on the box. Snowflake's design pulls those concerns apart into three layers that scale, fail and get billed independently. Almost every other Snowflake answer — caching, concurrency, cost, cloning — is a consequence of this split, which is why interviewers open here. ## Layer 1 — database storage When you load data, Snowflake does not keep your CSV or Parquet file. It reorganizes the rows into its own **compressed, columnar** internal format and writes them as immutable files (Snowflake calls them micro-partitions) into the cloud provider's object storage — Amazon S3, Azure Blob Storage, or Google Cloud Storage — inside the region the account lives in. You cannot open those files directly; the only access path is SQL through Snowflake. Alongside the data, Snowflake records per-file, per-column metadata (value ranges, counts, distinct-value information) that later drives pruning. Storage is billed on the average compressed bytes stored per month and is completely independent of whether any compute is running. A table you never query costs storage and nothing else. ## Layer 2 — query processing (virtual warehouses) A **virtual warehouse** is a cluster of compute nodes Snowflake provisions on your behalf. Each warehouse has its own CPU, memory and local SSD; it pulls the file chunks it needs from the storage layer and caches them on that local SSD for reuse. Warehouses are shared-nothing among themselves: warehouse `ETL_WH` and warehouse `BI_WH` share no compute state, so a heavy transformation cannot slow a dashboard running on the other warehouse. Any number of warehouses can read (and write) the same tables concurrently, because they all point at the same storage. Warehouses can be created, resized, suspended and resumed in seconds — nothing has to be redistributed, because they own no data. Credits are consumed only while a warehouse is running. ## Layer 3 — cloud services This is the brain, and it is a **multi-tenant** service Snowflake operates across accounts. It owns: - authentication, sessions, and role-based access control; - the catalog and all object metadata, including the per-file statistics used for pruning; - SQL parsing, optimization and compilation, and query dispatch to a warehouse; - transaction management and the ACID guarantees over the shared storage; - the **query result cache** and infrastructure management (provisioning warehouse nodes). Nothing here belongs to a single warehouse. That is exactly why metadata and cached results are shared account-wide, and why some statements complete with no warehouse running at all. Cloud services consumption is metered in credits, but Snowflake charges only the portion of daily cloud-services credits that exceeds 10% of that day's warehouse credits, so for normal workloads it rounds to nothing. ## What the architecture buys you - **Independent scaling.** Data growth never forces a compute purchase, and a bigger warehouse never forces a data reload. - **Workload isolation.** ELT, BI and data science each get their own warehouse over one copy of the data — no extracts, no marts kept in sync. - **Elasticity.** Resize or spin up a warehouse in seconds because there is no data to move. - **Cheap metadata operations.** Clones and time travel are pointer manipulations over immutable files, not copies. ## What it costs you - A cold warehouse pays object-store latency on its first scan until its local cache fills. - An idle running warehouse burns credits for nothing, so suspend policy is a real cost lever. - Compilation is centralized: a very large plan can spend meaningful time in the services layer before any data is read. - Nothing about this design targets single-row OLTP access; there are no user-managed indexes in the traditional sense. ## Interview framing The follow-up is usually "which layer does X live in?" Be able to place: pruning statistics (services metadata, describing storage), the query result cache (services, shared by all warehouses), the local data cache (inside one warehouse, lost on suspend), RBAC and the optimizer (services), the actual bytes (storage).

  • If you suspend a virtual warehouse, what is lost and what survives?
    The data itself and all metadata survive — they live in the storage and cloud services layers, not on the warehouse. What is lost is the warehouse's local SSD cache, so the next query after resume re-reads from object storage and runs colder. Query results already in the account-wide result cache still serve, because that cache belongs to cloud services.
  • Which layer decides how much data a query actually reads?
    Cloud services. The optimizer consults the per-file column metadata it keeps about the storage layer and eliminates files whose value ranges cannot satisfy the predicate, then hands the surviving file list to the warehouse. The warehouse only fetches and scans what it is told to.
  • Can two Snowflake accounts share the same storage layer?
    Not directly — each account's tables live in storage Snowflake manages for that account. Cross-account access is granted through Snowflake's secure sharing features, where the consumer queries the provider's data with the consumer's own compute; no copy is made and no files change hands.

Think of a public library: the shelves hold one copy of every book (storage), any number of reading rooms can work from those shelves at once (virtual warehouses), and the catalog, librarians and membership desk sit above both (cloud services).

saying these in an interview costs you the question

  • Says data is copied onto the warehouse nodes permanently
  • Claims each warehouse owns a slice of the table
  • Thinks storage costs stop when warehouses are suspended
  • Puts the query optimizer inside the virtual warehouse
  • Calls Snowflake pure shared-nothing like a classic MPP appliance

context