In Spark, what is the difference between a transformation and an action on a DataFrame?
answer
- typing the chain costs nothing
- one kind returns another DataFrame
- the other hands a result to the driver
- laziness is what lets Spark rewrite the plan
basics
~20 sA transformation such as filter or select only records intent and returns a new DataFrame; an action such as count, collect, show or write executes the accumulated plan on the cluster and returns a result outside Spark.
solid answer
~40 sSpark is **lazy**. Calling `filter`, `select`, `withColumn` or `join` builds up a description of the computation — a logical plan for DataFrames, a chain of parent RDDs for RDDs — and returns a new DataFrame without reading any data. Only an **action** asks for something that is not a DataFrame (`count()` returns a number, `collect()` returns rows to the driver, `write.parquet(...)` produces files), and that is what makes the driver optimize the plan, cut it into stages, and launch tasks on executors. Laziness is what lets the optimizer see the whole chain at once, so it can prune columns and push filters into the scan instead of executing step by step. The flip side: two actions on the same uncached DataFrame run the work twice.
code
python · 6 linesdf = spark.read.parquet("/data/orders") # no job (Parquet schema is in the footer)
big = df.filter(df.amount > 100) # transformation - plan only
cols = big.select("order_id", "amount") # transformation - plan only
n = cols.count() # ACTION -> job 1
rows = cols.take(5) # ACTION -> job 2, reads the source againgo deeper
Be ready to sort a list of calls into transformations and actions and to say that nothing runs until an action. Know count, collect, show and write as actions.
Explain what the driver does the moment an action fires: analyze, optimize, cut the plan into stages at shuffle boundaries, launch one task per partition. Name the reuse and error-timing consequences.
Show that you use laziness deliberately: reading explain() before running, caching only where a DataFrame is genuinely consumed twice, and recognising the calls that launch surprise jobs such as schema inference.
Be able to argue when laziness stops helping — pipelines whose plans grow so large that planning itself becomes the bottleneck, or long chains where materializing an intermediate table is better for cost, debuggability and restartability than one giant job.
## The two kinds of operation Everything you call on a Spark DataFrame or RDD falls into one of two categories. A **transformation** takes a DataFrame and returns another DataFrame: `select`, `filter`, `where`, `withColumn`, `join`, `union`, `distinct`, `orderBy`, `groupBy(...).agg(...)`, and on RDDs `map`, `flatMap`, `filter`, `reduceByKey`. A transformation does not touch a single row. For a DataFrame it appends a node to a *logical plan* — a tree describing what you want. For an RDD it creates a new RDD object holding a pointer to its parent plus the function to apply to each partition. An **action** asks for something that is *not* a DataFrame: `count()` (a number), `collect()` / `take(n)` / `first()` (rows materialized in the driver's JVM), `show()` (printed rows), `foreach` (a side effect on executors), `write.save(...)` / `write.parquet(...)` (files). That request is what makes Spark actually run. ## What happens the moment you call an action The driver resolves the plan against the catalog, lets the optimizer rewrite it, chooses a physical plan, and splits that plan into **stages** at every shuffle boundary. Each stage becomes a set of **tasks**, one per partition, dispatched to executors. Results stream back (or are written out), and the job ends. Nothing before that instant involved the cluster at all: `spark.read.parquet(p).filter(...).select(...)` on a petabyte-scale directory costs microseconds until you ask for a result. ## Why laziness is worth the confusion Because Spark sees the *whole* chain before it runs anything, it can rewrite it. If you filter after a join, the optimizer may push the predicate below the join and into the file scan so fewer rows are ever read. If you only `select` two of eighty columns, a columnar source reads only those two. If your action is `count()`, the payload columns may not be read at all. An eager, step-by-step engine would have already materialized the intermediate results and lost every one of those opportunities. ## The costs you must be able to name **No automatic reuse.** Each action re-executes the plan from the source. `df.count()` followed by `df.collect()` reads the input twice unless you `cache()`/`persist()` the DataFrame — Spark does not memoize results for you. (Shuffle map output written during one job can be reused by a later one, which is why the second run is sometimes faster, but that is an optimization, not a guarantee you should design around.) **Error timing.** For DataFrames, the plan is *analyzed* eagerly as you build it, so a misspelled column name raises `AnalysisException` on the line that referenced it. But data-dependent failures — a bad cast, a corrupt file, an out-of-memory task — only appear when the action runs, which is why a stack trace often points at `collect()` rather than at the transformation that caused it. RDDs give you neither check: a closure with a bug is silent until it executes on an executor. **Exceptions to laziness.** A few calls launch a job even though they look like reads or transformations. `spark.read.json(path)` and `spark.read.option("inferSchema", "true").csv(path)` scan the data to infer a schema before you have called anything. `Dataset.checkpoint()` is eager by default. Knowing these is a good sign in an interview. ## Same rule, both APIs RDDs and DataFrames share the laziness model; they differ in *what* is accumulated. An RDD chain is a lineage of opaque JVM functions — Spark knows only "apply this closure to each partition". A DataFrame chain is a plan of typed expressions Spark understands, which is exactly why the DataFrame API can be optimized and the RDD API mostly cannot. That is the single best reason to prefer DataFrames for anything expressible as columns. ## How to inspect without running `df.explain()` (or `df.explain(True)` in PySpark for the full plan set) prints the plan Spark would execute, without executing it. Reading that output is the fastest way to confirm that a filter really was pushed down or that a join really will broadcast — and it is a habit interviewers look for. ## What a good answer sounds like "Transformations are lazy and build a plan; actions trigger a job. Laziness lets Catalyst optimize the whole chain rather than each step, but it means repeated actions recompute unless I cache, and it shifts runtime errors to the action's stack trace."
- Which operations break the rule and launch a job even though they look like reads or transformations?Schema inference does: `spark.read.json(path)` infers by scanning the files, and CSV with `inferSchema` set to true makes an extra pass, both before any action. `Dataset.checkpoint()` is eager by default. Supplying an explicit schema removes the inference job.
- Why does calling count() and then collect() on the same DataFrame read the source twice?Because Spark does not memoize results. Each action re-plans and re-executes from the source. `cache()` or `persist()` marks the DataFrame for reuse, but even then the first action populates the cache and later ones read it — and cached blocks can be evicted, in which case Spark recomputes.
- If DataFrame transformations are lazy, why does a typo in a column name fail immediately?The logical plan is *analyzed* eagerly when you build the DataFrame, so name and type resolution against the catalog happens at definition time and throws `AnalysisException` on that line. Only execution is deferred, so data-dependent errors still surface at the action.
Transformations are writing a shopping list; the action is walking into the shop. Nothing is bought while you edit the list, and rewriting it beforehand is free.
saying these in an interview costs you the question
- Says filter() immediately scans the data on the cluster
- Thinks show() or count() is a transformation
- Believes Spark automatically caches results between actions
- Claims lazy means nothing runs unless you call cache()
- Confuses lazy evaluation with asynchronous or background execution