How would you choose between a managed warehouse's own storage and open table formats in your object storage?
answer
- who else needs to read these tables
- price per terabyte is the wrong comparison
- governance has to hold across every engine
- every knob you add someone must turn
basics
~20 sDecide on how many engines must read the data, how much governance and performance you need out of the box, and how much platform work you can staff. Warehouse-managed storage buys integration and predictability; open formats buy engine choice and exit options, paid for in operational ownership.
solid answer
~50 sFrame it around four axes. **Engine plurality:** if only SQL analytics reads the data, a managed warehouse is simpler; if batch, streaming, ML and SQL all need the same tables, open formats avoid copies that eventually disagree. **Total cost of the workload**, not storage price: warehouse-managed layout, clustering and caching usually cut bytes scanned, while open tables cut copies and let you retain history cheaply — model your real query mix rather than comparing per-terabyte rates. **Governance:** a warehouse gives one place for column and row policies, lineage and audit; across many engines on shared files you must build that consistently or accept gaps. **Exit cost:** with open formats the data stays yours and switching engines is a re-pointing exercise rather than an unload-and-reload migration. Most mature platforms split it: open tables as the curated shared layer, warehouse-managed storage for the hot, high-concurrency serving marts.
go deeper
Know that the choice is between data the engine owns in its own format and data you own in open formats readable by many engines, and that it is a trade-off rather than a ranking.
Be able to name the concrete consequences of each: copies and pipelines versus compaction and metadata upkeep, single-engine optimisation versus multi-engine access.
Argue it from a real workload — engine consumers, top queries by cost, governance requirements — and describe the hybrid where the shared curated layer is open and the hot serving marts are loaded.
Own the whole frame: engine plurality, modelled total cost, governance enforcement across engines, exit cost, and what the platform team can actually operate. State which axis dominates here and what evidence would change your mind.
## What is actually being decided The question is not "which product" but **who owns the physical layout of the curated data** — a single vendor's engine, or your object storage under an open table format that several engines can read. Everything else follows from that. ## Axis 1: how many engines must read this If SQL analytics is the only consumer, warehouse-managed storage is the low-friction answer: one system, one security model, one support contract, and the engine can optimise layout end to end. The moment a second consumer class appears — a distributed batch engine, a streaming job, a Python or ML workload, a second SQL engine after an acquisition — the managed-storage answer becomes "export a copy". Copies cost storage, they cost a pipeline to maintain, and above all they cost **agreement**: two copies of a table drift, and two teams arrive at a meeting with two numbers. Open table formats exist mostly to eliminate that class of copy. Verify per engine, though — multi-engine *read* support is broad, multi-engine *write* support is narrower, and "any engine can use it" is a claim to test rather than assume. ## Axis 2: cost, modelled on the real workload Comparing storage list prices is the wrong comparison. Warehouse-managed storage is usually more expensive per terabyte, but the engine's clustering, statistics and result caching can reduce bytes scanned by an order of magnitude on a well-shaped workload, and compute is normally the dominant line item. Open tables win on retention economics — years of history at raw object-storage rates — and on eliminating duplicate copies and their pipelines. The honest exercise is to take your actual top queries by spend, estimate bytes scanned under each layout, add copy and pipeline cost on one side and compaction, metadata maintenance and engineering time on the other, and compare. Do not assert a saving you have not modelled; the ratio is workload-specific. ## Axis 3: governance A managed warehouse gives you one enforcement point for column masking, row-level policies, auditing and lineage, and it is enforced no matter how the user connects. Shared open files invert this: the table format defines the data, but *who may read which column* is enforced by whatever catalog and storage permissions each engine honours. If three engines read the same prefix, you need a policy layer all three respect, or you accept that the weakest path defines your actual security posture. In a regulated environment this frequently decides the question on its own — and if it does not, budget explicitly for the catalog and permission work rather than discovering it later. ## Axis 4: lock-in and exit cost With open formats the bytes live in your bucket in a public format, so replacing an engine means re-pointing rather than unloading and reloading petabytes, and you can adopt a new engine for one workload without migrating everything. That optionality is real, but it is insurance, not a free good: you pay for it in operational ownership and usually in some performance. Weigh it against how likely you are to switch, and against the leverage it gives you at renewal time. ## Axis 5: what you can staff This is the axis teams skip. Open table storage makes you responsible for compaction of small files, expiry of old versions, orphan cleanup, catalog operation, and cross-engine upgrade compatibility. A managed warehouse does that housekeeping invisibly. If the platform team is two people supporting a hundred analysts, choosing the architecture with more knobs is choosing to have fewer of them turned. Be candid about this in the interview — it is the answer that separates a strategy from a preference. ## The shape most platforms converge on Raw data lands in object storage. The curated, shared layer is open table format — that is where multiple engines read, where history is retained, and where the authoritative definitions live. The hottest serving marts, behind interactive dashboards with high concurrency and tight latency, are loaded into warehouse-managed storage, where the engine's control of layout and caching pays for itself. The rule of thumb: **share on open formats, serve on managed storage.** ## How to make the decision defensible Write down the consumers and their engines, the top workloads by cost and by latency requirement, the regulatory constraints, and the team's capacity. Pick a small number of representative workloads and actually measure both layouts rather than arguing from architecture diagrams. Make the reversible choices first, and be explicit about which parts of the decision are hard to unwind — the curated layer's format is the sticky one; which engine queries it is comparatively cheap to change, and that asymmetry is itself an argument for keeping the curated layer open. ## What to avoid saying Avoid absolutes. "Open formats are always cheaper" ignores compute and engineering time. "Warehouses are always faster" ignores that the gap depends entirely on layout and workload. "Open formats mean no lock-in" ignores the catalog, the pipelines, the governance tooling and the SQL dialect you are still tied to. A principal-level answer names the axes, states which one dominates for the organisation in question, and says what evidence would change the call.
- Someone argues open table formats are simply cheaper. How do you test that claim?Model the whole workload, not storage. Take the top queries by spend, estimate bytes scanned under each layout, then add the pipeline and duplicate-storage cost avoided on one side against compaction, metadata maintenance and engineering time added on the other. Run a couple of representative workloads on both before committing.
- What is the part of this decision that is genuinely hard to reverse?The format and ownership of the curated layer. Which engine queries a table is comparatively easy to change; converting petabytes out of a proprietary managed store, and rebuilding every pipeline and permission that assumed it, is not. That asymmetry is the strongest structural argument for keeping the shared layer open.
- Does adopting open table formats actually eliminate vendor lock-in?No — it removes lock-in on the storage of the data, which is the expensive part, while leaving you tied to a catalog, a set of pipelines, a governance stack and a SQL dialect. It converts a data migration into an engine migration. That is a large improvement, not an escape.
Managed warehouse storage is a serviced apartment: everything works, the rules are the landlord's, and moving out means packing everything. Open formats are owning the building: full control and portability, and every repair is now yours.
saying these in an interview costs you the question
- Deciding on storage price per terabyte alone
- Claiming open formats remove vendor lock-in entirely
- Ignoring the compaction and metadata upkeep you take on
- Assuming every engine can write the tables, not just read them
- Treating governance as solved once the table format is chosen