In Spark, what is the difference between a job, a stage and a task?
answer
- nothing runs until you ask for a result
- cut wherever rows must move
- narrow steps get pipelined together
- one unit of work per partition
- ShuffleMapStage writes, ResultStage returns
basics
~10 sIn Spark, an action submits a job; the driver cuts that job into stages at shuffle boundaries; each stage runs one task per partition. Tasks are the smallest unit executors actually execute.
solid answer
~50 sA **job** is what a Spark action creates — `count()`, `collect()`, `write.parquet(...)` each submit one job, and transformations alone submit none. Inside the driver, the `DAGScheduler` walks the job's dependency graph backwards and cuts it into **stages** at every wide dependency, that is wherever data must be redistributed across the cluster; narrow operations like `filter`, `select` and `withColumn` are pipelined into the same stage. Each stage then produces one **task** per partition of its output. A `ShuffleMapStage`'s tasks write shuffle data for the next stage; the final `ResultStage`'s tasks compute the action's result. Tasks in a stage run identical code over different partitions, occupy one executor slot each, and are the unit Spark retries on failure. So the hierarchy is one action → one job → several stages → as many tasks per stage as that stage has partitions.
code
python · 6 linesfrom pyspark.sql.functions import to_date
df = spark.read.parquet("/data/events")
clean = df.filter(df.status == "ok").withColumn("day", to_date(df.ts))
daily = clean.groupBy("day").count()
daily.show() # first action here: one job, two stagesgo deeper
Be ready to name the three units in order and say what triggers each: an action makes a job, a shuffle makes a stage boundary, a partition makes a task. Know that transformations alone execute nothing.
Explain the mechanics: the DAGScheduler cuts at wide dependencies, narrow operators are pipelined into one stage, and task count equals output partitions. Distinguish ShuffleMapStage from ResultStage.
Show you can map a Spark UI stage page back onto this model — how many partitions, how many waves, which boundary came from which operator — and use it to explain why a job is slow.
Own the framing that stage count and task granularity are design choices with a cost curve: too few tasks wastes the cluster, too many drowns the driver in scheduling. Set team defaults accordingly.
## Why Spark has three units of work Spark is lazy. Transformations such as `filter`, `select`, `map` and `join` only build a description of a computation; nothing runs until you call an **action** — `count()`, `collect()`, `first()`, `show()`, `write.parquet(...)`. At that moment the driver holds a complete dependency graph and has to turn it into work it can hand to machines. It does that in three nested units: job, stage, task. ## Job — one per action An action submits a **job** to the `DAGScheduler` that runs inside the driver. A job is the whole computation needed to produce that one action's result, traced back to the sources. Two actions on the same DataFrame submit two jobs, and unless you cached something in between, the second job recomputes the shared work from scratch. The one-action-one-job rule has honest exceptions: a range-partitioned sort first runs a small sampling job to choose range boundaries, and some write and `show` paths submit more than one job. Jobs appear in the Spark UI's Jobs tab, each tagged with the call site that triggered it. ## Stage — a pipelineable segment, cut at shuffles The `DAGScheduler` walks the graph backwards from the action and cuts a new **stage** at every wide dependency: every point where a record's destination depends on its key, so rows must be redistributed across executors. `groupBy`, a join on a non-co-partitioned key, `distinct`, `repartition` and `orderBy` all create such a boundary. `filter`, `map`, `select` and `withColumn` do not. Everything between two boundaries is *pipelined*. Spark does not materialize an intermediate collection per operator; it composes them into one function applied record by record. That is why a chain of ten narrow transformations still runs as a single stage. There are exactly two kinds of stage. A `ShuffleMapStage`'s output is shuffle data written to local disk (or handed to an external shuffle service) for a downstream stage to fetch. The final `ResultStage` computes the action's result and returns it to the driver or writes it out. Dependent stages run in sequence; independent branches — for example both sides of a join — can run concurrently when slots are free. ## Task — one per partition Within a stage, Spark creates one **task** per partition of that stage's output. A task is the smallest schedulable unit: the same serialized plan fragment applied to one partition's data, shipped to an executor and run in one thread of that executor's JVM. A stage over 200 partitions is 200 tasks, all running identical code on different slices of data. Task count is a function of partitioning, not of cluster size. A 200-task stage on a cluster offering 40 concurrent slots simply runs in five waves. Tasks are also the unit of failure and retry: when a task dies, Spark re-runs *that task* on another executor (up to `spark.task.maxFailures`, default 4) rather than the whole stage. That is only safe because a task's input can be recomputed from lineage or re-read from existing shuffle files. ## How this shows up in the Spark UI The Jobs tab lists jobs with their triggering action and links to each job's stage DAG. The Stages tab shows per-stage task counts, durations, input size, shuffle read and shuffle write. Some stages are marked **skipped** — not an error, it means the shuffle output that stage would have produced already exists on disk from an earlier job, so Spark reused it instead of recomputing. ## Do not import vocabulary from other engines A MapReduce *task* is a whole mapper or reducer process. A Flink *subtask* is one long-lived parallel instance of an operator inside a continuously running dataflow. A Spark task is short-lived, does one partition's work inside one stage, and then its thread is reused for the next task. Interviewers notice when these are blurred together. ## The practical payoff Almost every Spark tuning conversation reduces to this hierarchy. "Too few tasks" means partitions are too coarse and slots sit idle. "Thousands of tiny tasks" means scheduling overhead dominates real work. "Eight stages where I expected two" means shuffles you did not intend. "One task in the stage never finishes" means one partition is far larger than the others. Being able to open a stage page and state how many partitions there are, how many waves they will run in, and where each boundary came from is exactly the skill being probed.
- Why does the Spark UI sometimes mark a stage as "skipped"?Because the shuffle output that stage would have produced is already on disk from an earlier job in the same application, so Spark reuses it instead of recomputing. It is a sign of successful reuse, not of a failure. You see it often when several actions share an upstream shuffle, or after caching a DataFrame.
- Can a single action ever submit more than one job?Yes. A range-partitioned sort or `orderBy` first runs a small sampling job to choose range boundaries, then the real job. Some write paths and `show()` on certain plans also submit extra jobs. So the Jobs tab can show more jobs than you wrote actions, which is normal rather than a bug.
- If a stage has 200 tasks and a cluster with 40 slots, how does it execute?In five waves of 40. Spark assigns tasks to free slots as they open, so the stage finishes after roughly five task durations plus scheduling overhead. Task count comes from partitioning, not from cluster size, so more executors reduce the number of waves but never change how many tasks exist.
An action is the order placed at a kitchen; stages are the courses that must be finished before the next can start; tasks are the identical portions each cook plates in parallel within one course.
saying these in an interview costs you the question
- Says every transformation starts a new job
- Claims a stage boundary appears at each operator
- Confuses a Spark task with a MapReduce mapper or reducer
- Thinks executor count determines how many tasks a stage has
- Believes dependent stages run in parallel with each other