When does per-byte-scanned billing beat paying by the second for a provisioned warehouse cluster?
answer
- what are you actually renting: work or a machine
- idle time is the pivot
- how busy would the cluster really be
- a full scan costs the same either way under one model
- cache reuse only pays under one model
basics
~20 sPer-byte-scanned billing wins on spiky, low-duty-cycle workloads whose queries prune well, because you pay nothing between queries. Provisioned compute wins when a cluster stays busy, when queries repeatedly scan the same large data, and when caching amortizes across many statements.
solid answer
~50 sSeparation of storage and compute makes both models possible, and the deciding variable is duty cycle. Under per-byte-scanned pricing you pay for data read regardless of how long it took or how many machines ran, so idle time is genuinely free and a well-pruned query is cheap — but a full scan costs the same whether it feeds a `COUNT(*)` or a complex report, and a dashboard refreshing every minute over a large table is expensive forever. Under provisioned billing you pay cluster size times wall-clock time, so a cluster that is busy most of the hours it runs is very cheap per query, repeated scans get nearly free once cached locally, and cost is bounded by the hardware you chose — but idle minutes burn money unless the cluster suspends. To decide, measure two numbers over a representative week: total bytes scanned, and the fraction of time a cluster would actually be busy. Then compare bytes × per-TB rate against size × hours.
code
text · 8 lines-- substitute your own contracted rates; do not assume list prices
per_byte_month = bytes_scanned_per_month_TB * rate_per_TB
provisioned_month = cluster_units * hours_cluster_is_up * rate_per_unit_hour
-- workload A: 400 TB scanned/month, cluster would be busy 3% of the month
-- workload B: 40 TB scanned/month, cluster would be busy 70% of the month
-- A favours capacity pricing only if pruning cannot cut the 400 TB
-- B almost always favours provisioned: high duty cycle + cache reusego deeper
Know that one model charges for the data a query reads and the other charges for the time a cluster is up, and that idle clusters still cost money.
Explain what each model does and does not charge for — query complexity, concurrency, runtime, bytes — and why a final LIMIT does not shrink a scan.
Show you would measure bytes scanned and duty cycle before choosing, and name the workload patterns that blow up each model in production.
Own the portfolio decision: which workloads sit on committed capacity, which stay on-demand, how much to commit against forecast risk, and how billing shape steers your team's tuning habits.
## Why there are two models at all Once compute is stateless and can be started and stopped freely, a vendor can meter it in two quite different ways. Meter the **work** — bytes read from storage — and the customer never sees a machine. Meter the **capacity** — cluster size multiplied by the seconds it runs — and the customer chooses the machine and owns its utilization. Both exist because different workloads have opposite shapes, and mature platforms increasingly offer both plus a committed-capacity option in between. ## What each model actually charges for **Per-byte-scanned.** The unit is data read by the query, typically compressed bytes for the columns and files the engine actually touched. Consequences that surprise people: cost is independent of how long the query runs and how much CPU it burned, so an expensive multi-way join over a small scan is cheap and a trivial aggregate over a huge scan is not. Cost is also independent of concurrency — a hundred simultaneous queries cost the sum of their scans, not a queue. And a final `LIMIT` usually saves nothing, because the rows must be read before they can be limited. **Provisioned by the second.** The unit is cluster size times uptime. Consequences: query complexity and bytes read are free at the margin, so the incentive is to pack as much work as possible into the hours the cluster is up. Caching compounds this — the second query over the same data is faster and therefore literally cheaper. But an idle running cluster bills at full rate, which is why idle-suspend behaviour becomes a first-order cost control. ## The comparison you actually do Instrument a representative week and get two figures per workload: total bytes scanned, and the wall-clock time a suitably sized cluster would be executing. Then: ``` per-byte cost = bytes_scanned_per_month x rate_per_TB provisioned = cluster_units x hours_running x rate_per_unit_hour ``` Substitute your own contracted rates; the shape of the answer, not any published price, is what matters. The break-even is a duty-cycle line. Below it — a handful of analysts, an hourly job, an unpredictable data-science workload — per-byte usually wins because you buy nothing when nobody queries. Above it — an ELT pipeline running most of the night, BI serving hundreds of users all day — provisioned usually wins, and by a wide margin once cache reuse is counted. ## The patterns that break each model Per-byte billing is punished by: - **Unpruned repetition.** A dashboard tile refreshing every minute against an unpartitioned fact table scans the whole thing every minute. This is the classic surprise invoice. - **`SELECT *` in a columnar store.** Reading every column when five are needed multiplies the bill directly. - **Cheap-looking probes.** `SELECT ... LIMIT 10` on a wide table reads the columns it projects, and previewing a table repeatedly is not free. - **Growth.** Scan volume tracks table size, so flat query volume over tripling data means a tripling bill unless filters prune. Provisioned billing is punished by: - **Idle clusters.** A cluster left running overnight for a job that finished at 01:00 is pure waste. - **Cluster sprawl.** Many small clusters each idling below their suspend threshold cost more in aggregate than one shared one. - **Over-sizing for the worst query.** Buying the size the nightly rebuild needs and leaving it up for interactive work all day. - **Cold starts fighting the suspend timer.** Suspending too aggressively on a bursty workload lengthens every query and can raise total compute-seconds. ## Hybrids and the honest answer Most platforms now let you mix: reserve a baseline of capacity at a discount and let overflow bill on demand, or run production pipelines on provisioned compute while ad-hoc exploration bills per byte. That combination is usually the right recommendation in an interview, because it matches the two workload shapes rather than forcing one to pretend to be the other. Say explicitly what you would measure before choosing, and note that reservations trade flexibility for rate — under-used committed capacity is a real way to lose the savings you bought. One more point worth making: the two models create opposite optimization incentives, and that shapes engineering effort. Per-byte pricing pays you directly for partitioning, clustering and narrow projections, and pays you nothing for making a query algorithmically faster. Provisioned pricing pays you for concurrency packing, cache warmth and shorter wall-clock runtime, and pays you nothing for scanning fewer bytes if the cluster was going to be up anyway. Teams inherit the habits their billing model rewards, and a migration between models frequently invalidates a year of tuning intuition.
- How does per-byte-scanned billing change how you design tables?It converts layout into money. Partitioning and clustering on the columns your filters use, keeping wide tables narrow or splitting rarely-read columns out, and avoiding `SELECT *` all reduce the invoice line by line. Under provisioned billing the same changes only reduce runtime, which matters but does not show up as a distinct charge.
- What is the risk of buying committed or reserved capacity to lower the rate?You pay for it whether or not you use it, so a workload that shrinks, moves, or turns out to be spikier than forecast leaves you paying for idle capacity at a discounted rate that is still worse than paying nothing. Commit to the floor of your demand, not the peak.
- Under per-byte billing, why does adding LIMIT to a report rarely reduce cost?The engine must read the columns and files needed to determine which rows qualify before it can limit the output, and any sort or aggregation forces a full read anyway. Limiting shrinks the result set delivered to the client, not the data scanned to produce it.
saying these in an interview costs you the question
- Says per-byte billing is always cheaper because there is no idle cost
- Thinks a long-running complex query costs more under per-byte pricing
- Believes adding LIMIT cuts the bytes billed
- Ignores duty cycle and compares only headline rates
- Assumes reserved capacity is free money with no downside