What is the difference between ETL and ELT in a data pipeline?
answer
- same three steps, different order
- where does the T actually run
- outside the target, or inside it
- one pattern keeps the source-shaped rows
basics
~20 sETL transforms data in a separate processing tier before writing it to the target. ELT loads source-shaped data into the target first and transforms it there, using the target's own compute, keeping the raw input queryable.
solid answer
~50 sBoth patterns do the same three operations — extract, load, transform — and differ only in where the transform runs. In **ETL**, data is pulled into a separate processing tier (an ETL server, a mapping tool's engine, a processing cluster), reshaped there, and only the finished result is written to the target; the target never sees source-shaped rows. In **ELT**, extraction lands data in the target with as little change as possible and the reshaping runs afterwards as statements the target executes itself — usually SQL over tables already in the warehouse. The practical consequence is not stylistic: ELT keeps an untouched copy of the input, so a bug in the logic is fixed by re-running SQL, while an ETL pipeline that discarded its input has to go back to the source. ELT also puts transformation cost on the warehouse bill rather than on a cluster you operate.
code
text · 2 linesETL: source -> [processing tier: clean, join, aggregate] -> target (curated only)
ELT: source -> target (raw) -> [target runs SQL] -> target (curated)go deeper
Be able to spell out the three operations and say which one moves. Recall that ELT lands source-shaped data in the target first and reshapes it there, and that ETL finishes the reshaping before anything is written.
Explain the mechanics behind the swap: which system's compute runs the logic, what language it is written in, and why keeping the landed input makes a logic fix a re-run rather than a re-extraction.
Show that you treat this as an architecture decision, not a fashion. Name the cases where transform-before-load is still mandatory — restricted fields, unparseable formats, volume you refuse to move — and describe the hybrid you would actually run.
Own the consequences across the platform: who pays for transformation compute, what skills the choice demands, and how retaining raw interacts with retention, deletion and audit obligations.
## The three operations Every analytics pipeline does three things. **Extract** reads data out of a source system — an application database, a SaaS API, a partner's file drop. **Load** writes it into the target analytical store. **Transform** reshapes it: fixing types, deduplicating, joining across sources, applying business rules, aggregating. ETL and ELT are the same three operations. They differ only in whether the transform happens *before* the load or *after* it. ## ETL: transform between the systems In ETL, data is extracted into a separate processing tier, transformed there, and only the finished result is written to the target. That processing tier is real infrastructure someone owns: a dedicated server running a vendor's engine, a graphical mapping tool, a distributed processing cluster, or hand-written code on a scheduled box. The target sees only rows that already conform to the intended model. That has consequences. The target stores curated output only, so it stays small and tidy. The source-shaped data never lands anywhere queryable — often it exists only as bytes in flight. The transformation is written in whatever language the processing tier speaks. And the pipeline's throughput is the capacity of that tier, which has to be sized, patched, monitored and paid for whether or not it is busy. ## ELT: transform inside the target In ELT, extraction writes source-shaped data straight into the target with minimal change, and all the reshaping runs afterwards as statements the target executes — typically SQL reading tables and writing tables inside the warehouse or lakehouse. Business logic becomes queries, and those queries run on the same engine that serves analysts. The consequences invert. Raw data is in the platform and can be inspected and re-read. Transformations are set-based operations over columns rather than procedural row-at-a-time processing. There is no separate cluster to operate. The target's bill now covers transformation as well as reporting. And because the untouched input is retained, any downstream table can be rebuilt by re-running the logic against data that is still there. ## What actually moves The shorthand is "the L and the T swap places", but the substantive change is that *the compute doing the transformation* moves — from a system the data team operates into the system that already holds the data. Everything else follows from that one relocation: who pays, what language the logic is written in, whether the raw survives, and how much data has to cross a network. ## Why the raw copy is the decisive difference In ETL, the transformed output is usually the only surviving representation of that day's data. If the business rule was wrong, the correct history may no longer be obtainable: operational sources update rows in place, purge old records, and rate-limit or bill for large re-reads. In ELT, the landed extract is still sitting there, so correcting logic is a re-run rather than a re-extraction. Interviewers care much more about this property than about the letters. ## "ELT" does not mean "nothing happens before the load" Even an ELT pipeline does technical work in flight: decompressing files, decoding character sets, parsing a nested payload into a tabular shape, assigning types, partitioning by ingestion time, and frequently masking or tokenising restricted fields that policy forbids from landing in raw form. What moves to the target is the *business* transformation — cross-source joins, deduplication, derived metrics, conformed naming. A candidate who says "ELT means we don't transform" has misread the pattern. ## When each pattern still fits ELT is the modern default when the target is an elastic columnar analytical engine that can absorb raw volume cheaply and scan it fast. ETL is still the right answer when the target cannot transform at reasonable cost (a small operational database, a fixed-capacity appliance); when the raw data legally may not land in the target; when raw volume is enormous relative to the useful slice, so moving all of it is the dominant cost; or when the work is not expressible in the target's query language at all — parsing binary formats, calling an external service per record, running a model. Most real stacks are hybrids: an EL connector lands raw, in-warehouse SQL builds the models, and a handful of jobs that genuinely need a general-purpose engine run outside and write their results back in. ## Neighbouring vocabulary **EL** on its own describes a connector that only lands data and leaves modelling to someone else — the first half of ELT. **Reverse ETL** is a different pattern entirely: pushing modelled warehouse tables back out into operational SaaS systems so sales and support tools can use them. Neither is a variant of the ETL-versus-ELT choice.
- Does ELT mean nothing is transformed before the data lands?No. Technical work still happens in flight: decompression, character-set decoding, parsing nested payloads into a tabular shape, assigning types, partitioning by ingestion time, and masking fields that policy forbids storing raw. What moves into the target is the business transformation — cross-source joins, deduplication, derived metrics and conformed naming.
- Where does the transformation code physically execute in each pattern?In ETL, on a processing tier your team provisions and operates between source and target — its CPU and memory are the pipeline's ceiling. In ELT, inside the target analytical engine, which parallelises the statement across its own workers. The orchestrator in either case issues the instruction; it should not be doing the data work itself.
- Is reverse ETL just ETL run backwards?No. Reverse ETL takes tables already modelled in the warehouse and pushes them out to operational SaaS systems — a CRM, a support tool, an ad platform — so business users act on them in the tools they live in. It is a delivery pattern for finished data, not a choice about where transformation happens.
ETL is cooking at home and bringing a finished dish to the party; ELT is bringing the raw ingredients and cooking in the host's much larger kitchen — where anyone can also see what went in.
saying these in an interview costs you the question
- Says ELT means the data is never transformed
- Claims ETL is obsolete and never correct today
- Thinks the difference is which vendor's tool you bought
- Confuses ELT with reverse ETL into SaaS systems
- Believes ELT requires a data lake rather than a warehouse