Why did cheap warehouse compute make ELT the default over ETL?
answer
- the economics changed, not the theory
- a rented query versus a standing cluster
- storage got cheap, scans got fast
- the transform is now SQL in a repository
basics
~20 sElastic columnar warehouses over cheap object storage made loading raw and transforming with SQL cheaper and faster than provisioning a separate ETL cluster. Transformation became rented per query, written in SQL, and rerunnable against retained raw data.
solid answer
~50 sThe theory did not change; the economics did. Analytical engines that separate storage from compute made it cheap to store source-shaped data and fast to scan it, and they rent transformation capacity per query instead of demanding a cluster sized for the annual peak. Once the warehouse could out-run a bespoke processing tier on ordinary joins and aggregations, keeping a second engine alive purely to reshape data stopped paying for itself. Three secondary effects locked the default in: extraction became a commodity — connectors that only land data — so teams stopped writing bespoke pipelines; transformation became SQL, which far more people can read, review and version-control; and retaining the raw input made logic fixes a re-run rather than a re-extraction. The counterweight is that transformation cost moved onto a per-query bill that grows quietly with every new model.
go deeper
Recall that ELT became practical because storing raw data got cheap and analytical engines got fast enough to transform it in place, so a separate transformation server stopped being necessary.
Explain the mechanics: elastic per-query compute versus a peak-sized standing cluster, columnar scans, commodity connectors, and logic that lives as reviewable SQL in version control.
Show the counterweight. Describe how a per-query transformation bill grows with model count and rebuild frequency, and what you put in place — attribution, incremental builds, review — to keep it honest.
Own the platform-level tradeoff: the pattern removed a cost governor along with a constraint, so budget discipline now has to come from ownership, showback and design review rather than from procurement.
## The change was economic, not conceptual ETL was not a mistake. It was the correct design for the targets that existed when it became standard. A warehouse ran on a fixed appliance whose capacity you bought years in advance; storage on it was expensive and rationed; and its query engine was tuned to serve reports, not to grind through raw source data. Under those constraints, landing raw data in the warehouse was unaffordable and transforming inside it competed directly with the reports. So teams bought a second engine to do the reshaping and sent the warehouse only the finished, small, modelled result. Every one of those constraints was relaxed by the cloud analytical engines that followed, and the default flipped. ## Storage stopped rationing the raw copy When table data sits on commodity object storage billed per gigabyte-month, keeping a full source-shaped copy of everything you ingest goes from a budget line someone has to defend to a rounding error against the compute bill. That single change is what makes the L-before-T possible at all: ELT's whole premise is that landing everything raw is affordable. ## Compute became elastic and set-based Two properties matter. First, **elasticity**: transformation capacity is rented while a statement runs rather than provisioned for the yearly peak and idle the rest of the time. A nightly rebuild that needs a large cluster for forty minutes pays for forty minutes. A dedicated ETL tier sized for the same peak is paid for continuously. Second, **columnar, massively parallel execution**: reading three columns of a wide table touches only those columns' data, and the engine splits the scan across many workers automatically. Ordinary analytical work — filter, join, group, window — is exactly what these engines are built for, and they typically beat a general-purpose processing tier at it without anyone tuning partitions or memory. ## Extraction became a commodity Once transformation moved into the target, the extract step no longer had to know anything about the model. It only has to land the source faithfully. That is a solved, buyable problem: managed connectors, log-based change capture, and file drops all just deliver rows. Teams stopped spending their scarcest engineers on plumbing whose only job was to move bytes, and the *E* and *L* collapsed into a single commodity step that someone else operates. ## Transformation became reviewable SQL This is underrated in interviews. When the logic is SQL executed by the warehouse, it lives in text files in a repository. It can be diffed in a pull request, tested, documented, and rebuilt from scratch by anyone with warehouse access. Compare that to a proprietary graphical mapping whose meaning lives inside a vendor tool and whose change history is whatever the tool records. Moving the transform into the target also moved it into normal software engineering practice, and it widened the pool of people who can safely change a pipeline from "the data engineering team" to "anyone who writes SQL". ## Reprocessing became cheap Because the input survives, a bug in a business rule is corrected by editing a statement and re-running it over data that is still sitting in the platform. In an ETL pipeline that discarded its input, the same fix requires going back to sources that update rows in place, purge history, throttle large reads, or bill for them. Teams that have lived through one bad restatement rarely give up the raw copy again. ## What still has to happen before the load ELT-first does not mean everything lands untouched. Restricted fields — payment credentials, health identifiers, anything a policy or regulation forbids storing in raw form — are masked, tokenised or dropped in flight, because "we will clean it up after it lands" is not a defence once it has landed. Formats the target cannot parse are converted first. And when the raw volume dwarfs the useful slice, filtering before the move is the cheaper design regardless of fashion. ## The bill that came with the default The honest counterweight: ELT converts a fixed, visible cluster cost into a variable, diffuse per-query cost. A separate ETL tier at least made you decide, once, how much transformation capacity the company was buying. In an ELT stack, every new model quietly adds recurring compute, nobody's individual query looks expensive, and the total grows with headcount. That is why cost attribution per model and disciplined incremental rebuilds show up so quickly on mature ELT platforms — the pattern removed a constraint that had also been acting as a governor.
- What must still be transformed before the load even in an ELT-first stack?Anything policy forbids storing raw — payment credentials, health identifiers, secrets — must be masked, tokenised or dropped in flight, since landing it first is already the violation. Formats the target cannot parse are converted before load, and when raw volume dwarfs the useful slice, filtering before the move is simply cheaper than moving it all.
- What did teams give up by adopting ELT?A fixed, visible transformation budget. A dedicated ETL tier forced one explicit decision about how much capacity the company bought; ELT spreads that cost across thousands of individually cheap queries that nobody reviews. The result is a bill that grows with the number of models and the rebuild frequency, which is why cost attribution and incremental rebuilds become mandatory disciplines.
- Does ELT depend on the target being a cloud warehouse?It depends on the target having elastic, cheap compute over cheap storage and a capable set-based query language — which cloud warehouses and lakehouse engines both provide. Against a fixed-capacity appliance or a small operational database, landing raw and transforming in place competes with the queries that matter, and transform-before-load remains the better design.
saying these in an interview costs you the question
- Says ELT is better because it is newer
- Claims warehouse compute is effectively free
- Ignores that transformation cost moved onto the query bill
- Says every field can be cleaned after it lands
- Treats the choice as independent of what the target can do