skip to content

For a company-wide analytical warehouse, how do you choose between one compute pool per team and shared pools with priority classes?

level: principalimportance: should knowfreq 40%

answer

  1. isolation is bought with utilisation
  2. teams are not a homogeneous workload
  3. the org chart is the wrong axis
  4. idle pools and cold caches cost real money
  5. make each team see its own bill

basics

~20 s

Split by workload class and SLA, not by org chart. Give guaranteed isolation only where a latency or freshness commitment justifies its idle cost; let everything else share pools with priorities, chargeback by usage, and revisit as workloads change.

solid answer

~50 s

The two poles trade **isolation** against **utilisation**. A pool per team gives perfect blast-radius containment and trivially attributable cost, but every pool idles separately, caches are duplicated instead of shared, and twenty pools mean twenty things to size and monitor. Shared pools with priority classes keep utilisation high and caches warm, but one team's mistake becomes everyone's incident, and cost attribution has to be reconstructed from query history. In practice the winning split is by **workload class and SLA** rather than by org chart: an interactive lane for dashboards, a batch lane for pipelines, an ad-hoc lane for exploration, plus a dedicated pool for the few workloads with a hard external commitment — a customer-facing product, or a regulatory deadline. Attach showback or chargeback from query-level usage so teams see their consumption whichever pool they use, and review the layout as workloads grow, because today's shared tenant is next year's dedicated one.

code

text · 14 lines
text
# split by workload class, with dedicated pools by exception

shared pools
  interactive : many teams, small grants, high concurrency, 5m cap
  batch       : many teams, large grants, low concurrency, 4h cap
  adhoc       : many teams, strict guardrails, lowest priority

dedicated pools (by policy, reviewed quarterly)
  customer_api : external latency SLA
  finance_close: regulatory deadline, month-end only, auto-suspend

cross-cutting
  per-query usage tagged by team -> showback in every pool
  spend ceiling per pool -> alert, then suspend

go deeper

for a junior

Know that a company-wide warehouse usually runs several compute pools rather than one, and that the split exists so heavy jobs cannot slow interactive users.

for a middle

Be able to state the core trade — isolation versus utilisation — and name the concrete costs of many small pools: idle compute, duplicated caches, more objects to size and monitor.

for a senior

Argue the split by workload class with evidence: which workloads measurably interfere, which have hard SLAs, and what guardrails and reservations you would apply inside each shared pool.

for a principal

Own the policy and its economics — who is entitled to dedicated compute and on what criteria, how usage is attributed and charged, what spend ceiling applies per pool, and the review cadence that keeps the layout matched to the workload rather than to the org chart.

## What the two poles actually buy **Pool per team** - *Isolation*: complete. One team's cross join cannot touch another team's dashboards. The blast radius of any mistake is one team. - *Attribution*: trivial. The pool's bill is the team's bill; no allocation model to argue about. - *Autonomy*: each team sizes and schedules its own compute without a central queue. - *Cost*: every pool has its own idle time and its own warm-up. Twenty pools at 30% utilisation cost far more than a handful at 80%. - *Cache fragmentation*: shared dimension tables get cached separately in every pool, so the same bytes are fetched from remote storage many times over. - *Operations*: twenty things to size, monitor, alert on and right-size, which in practice means most of them are never tuned at all. **Shared pools with priority classes** - *Utilisation*: high. One team's trough absorbs another's peak, and caches stay warm across users of the same tables. - *Fewer objects to manage*: a small number of pools that actually get attention. - *Interference*: real. Priorities govern admission but not the resources an in-flight query already holds, so a large query still degrades neighbours during execution. - *Attribution*: requires reconstructing per-team usage from query history, and someone has to own and defend that model. - *Blast radius*: organisation-wide. A single misconfigured refresh loop can affect everyone. ## The better axis: workload class, not org chart Teams are not a natural isolation boundary — a single team submits sub-second dashboard queries *and* four-hour rebuilds, so a per-team pool must be sized for the worst case and then does a poor job on the interactive one. Workload class is the axis where requirements are actually homogeneous: - **Interactive / BI**: many small queries, tight latency SLA, high concurrency, small grants, short timeouts. - **Batch / pipeline**: few large queries, deadline-based SLA ("fresh by 07:00") rather than latency, big grants, low concurrency, long timeouts. - **Ad-hoc / exploration**: unpredictable, occasionally pathological, deserves strict guardrails and the lowest priority. - **Customer-facing / external commitment**: anything whose latency is promised to someone outside the company. This is where a genuinely dedicated pool earns its idle cost. Within a class, teams share, because their queries look alike and their caches overlap. ## Deciding who gets dedicated compute Make it an explicit, defensible policy rather than a first-come-first-served negotiation. Reasonable criteria: 1. **An external or contractual SLA.** If missing a latency or freshness target has consequences outside the company, buy the isolation. 2. **Demonstrated interference.** Measured evidence that this workload degrades others, or is degraded by them, after priorities and caps have been tried. 3. **Scale that dominates.** A workload consuming a large share of a shared pool is effectively dedicated already; separating it makes cost visible and sizing honest. 4. **A compliance or residency requirement** that mandates separation. And criteria that should *not* qualify: seniority of the requester, a preference for "our own environment", or an unmeasured suspicion of interference. ## Cost visibility is the real lever The organisational failure mode is invisible cost: on a shared pool nobody's dashboard has a price, so nobody optimises. Attribute usage per query — by role, by service account, by tag — and publish it per team regardless of pool layout. Showback (visibility) changes behaviour surprisingly often; chargeback (actual budget transfer) changes it reliably but invites gaming, such as teams hoarding a dedicated pool to protect their allocation. Pair either with a **budget ceiling per pool** so a defect produces a suspended pool and an alert rather than an unbounded bill. ## Keeping it alive Whatever layout you choose is a snapshot of today's workloads and will decay. Build in a review: - Track per-pool utilisation, queue wait and spend. A pool consistently idle should be merged; a pool consistently queueing should be split, scaled, or have its worst tenant moved out. - Track guardrail trips per team — a rising rate is a workload-design problem surfacing as a platform problem. - Re-run the dedicated-pool criteria periodically, in both directions. Workloads graduate into dedicated pools as they grow, and should graduate back out when they shrink or when the product they served is retired. ## The answer an interviewer wants Not a doctrine but a method: start shared and cheap, isolate on measured evidence and on external commitments, split by workload class rather than by team, make cost visible everywhere, cap spend per pool, and schedule a review so the layout tracks the workload instead of the org chart.

  • What is the hidden cost of giving twenty teams twenty dedicated pools?
    Idle compute and fragmented caches. Each pool has its own trough, so aggregate utilisation collapses compared with pooling the same demand, and shared dimension tables are cached separately in every pool instead of once. You also create twenty objects to size, monitor and right-size, which in practice means most are never tuned and quietly run oversized. Auto-suspend reduces the idle cost but does not fix the cache fragmentation or the operational load.
  • How do you make cost visible when several teams share one pool?
    Tag usage at the query level — by role, service account, or an explicit workload tag — and roll it up per team from query history. Publish it as showback first; visibility alone changes behaviour more than people expect. Move to chargeback only when you are prepared to defend the allocation model, since real budget transfer invites gaming, such as teams demanding dedicated pools purely to fence off their allocation.
  • What evidence would move a workload from a shared pool to a dedicated one?
    Measured interference that survived the cheaper fixes: the workload's p95 queue wait or execution time degrades materially when specific neighbours run, after priorities, memory caps and reserved slots have already been applied. An external latency or freshness commitment qualifies on its own. So does sheer scale — a tenant consuming most of a shared pool is effectively dedicated already, and separating it makes both sizing and cost honest.

An office can give every team a private meeting room or share a well-scheduled pool of rooms. Private rooms guarantee availability and sit empty most of the day; shared rooms are efficient until two teams need one at once. Most buildings do both, split by how the room is used rather than by who works there.

saying these in an interview costs you the question

  • Splits compute along the org chart rather than by workload class
  • Treats a dedicated pool as free because storage is shared
  • Assumes priorities alone can meet a hard external latency SLA
  • Leaves per-team cost invisible on shared pools
  • Sets the layout once and never revisits it as workloads grow

context