In Spark, what does spark.speculation do, and when does it make a job worse?
answer
- a race between two copies of one task
- off unless you turn it on
- needs most of the stage finished first
- compare against the median duration
- useless when the partition itself is fat
basics
~20 sIt relaunches a duplicate of any task running far longer than its stage's median, and keeps whichever attempt finishes first. It is off by default and wastes slots when the slowness is skew rather than a bad node.
solid answer
~50 s`spark.speculation` is off by default. When enabled, the driver periodically inspects each stage's running tasks: once at least `spark.speculation.quantile` (default 0.75) of the stage's tasks have finished, any task still running longer than `spark.speculation.multiplier` (default 1.5) times the median completed duration is re-launched as a duplicate attempt on a different executor. The first attempt to finish wins and the other is killed. The check runs every `spark.speculation.interval` (default 100ms). It is a good defence against **environmental** stragglers — a failing disk, a noisy neighbour, a throttled node. It is a bad defence against **data** stragglers: if one task is slow because its partition holds ten times the rows, the duplicate is equally slow, so you burn a second slot for nothing. It is also unsafe when tasks have non-idempotent external side effects, since two attempts genuinely run.
code
properties · 4 linesspark.speculation=true
spark.speculation.interval=100ms
spark.speculation.quantile=0.75
spark.speculation.multiplier=1.5go deeper
Know that Spark can launch a duplicate copy of a slow task and keep whichever finishes first, and that this behaviour is off unless someone enables it.
Explain the trigger mechanics: a quantile of completed tasks establishes a median, and a multiplier above it marks a candidate, with the check repeating on a short interval.
Decide from evidence — compare the straggler's record counts to the median before enabling it — and articulate the side-effect risk for jobs writing outside Spark's commit protocol.
Own it as a cluster-wide policy question: which fleets get it by default given spot and node variance, which workloads must be excluded for idempotency, and what it costs in wasted slots.
## What it does Speculative execution is Spark's answer to the straggler problem: within a stage of otherwise identical tasks, one task takes far longer than the rest and the whole stage waits for it. With `spark.speculation=true`, the driver runs a periodic check — every `spark.speculation.interval`, default 100 milliseconds — over each active stage. The check needs a baseline, so it only arms once a fraction of the stage's tasks have completed: `spark.speculation.quantile`, default 0.75. From those completed tasks Spark takes the median duration. Any task still running whose elapsed time exceeds `spark.speculation.multiplier` (default 1.5) times that median becomes a speculation candidate, and the scheduler launches a second attempt of the same task on a different executor if a slot is available. Both attempts race; the first to succeed has its output committed and the other is killed. Because the trigger depends on 75% of the stage having finished, speculation does nothing for a stage with very few tasks, and nothing at all for a straggler that appears early in a long stage. ## When it earns its keep Speculation is designed for **environmental** variance, which is real on large or heterogeneous clusters: one node with a degraded disk, a machine whose CPU is being shared with a noisy co-tenant, a container throttled by cgroup limits, a NIC with retransmits, an object-store request that hangs. In all these cases the same task on a different machine simply runs at normal speed, so the duplicate finishes quickly and the stage stops waiting. On big shared clusters this is a meaningful tail-latency win, which is why many platform teams turn it on globally. ## When it makes things worse **Data skew.** If a task is slow because its partition contains far more rows — a hot key, a null key catching everything, an exploded array — then the duplicate attempt processes exactly the same oversized partition and is exactly as slow. You have doubled the work, occupied a slot another task could have used, and gained nothing. Worse, the duplicate can itself trip the threshold repeatedly on a saturated cluster. Skew needs a different fix entirely; speculation is a masking mechanism that will not mask it. **A saturated cluster.** Speculative copies compete for the same slots as real work. When every slot is busy, launching duplicates delays pending first attempts, so total throughput drops even though individual stages look better instrumented. **Non-idempotent side effects.** Speculation means two attempts of the same task genuinely execute. Spark's file commit protocol handles this for its own writes — only the winning attempt's output is committed and the loser's staged output is discarded — but it can say nothing about arbitrary code in a `foreachPartition` that posts to an HTTP API, increments a counter in an external store, or appends to a non-transactional sink. Those effects happen twice. This is the reason to think before enabling it globally on a platform that runs arbitrary user jobs. **Legitimately long tasks.** Some stages have a genuine bimodal distribution — a few partitions really do carry more work by design. Speculating on them is pure waste, and the noise in the UI (killed attempts, duplicate task rows) makes real diagnosis harder. ## Diagnosing before enabling The decision hinges on one question: is the slow task slow because of *where* it runs or because of *what it holds*? The Spark UI answers it directly. Open the stage's task table and compare the straggler's shuffle-read and input record counts against the median task's. If the record count is comparable and only the duration is an outlier, it is environmental and speculation will help; if the straggler read ten times the records, it is data and speculation will not. It is also worth checking whether the straggler always lands on the same host across runs — a stable answer points at a specific bad node, which is better fixed by excluding that node than by racing it. ## Practical stance A reasonable production default is: enable it on large, heterogeneous or spot-heavy clusters where node-level variance is common, and where jobs write through Spark's own commit protocol; keep it off for jobs with external side effects; and never treat it as a skew remedy. Tightening the multiplier makes it more aggressive and more wasteful; loosening it makes it fire too late to help.
- How do you tell whether a straggler is environmental or caused by skew?Compare the straggler's input and shuffle-read record counts with the stage's median task in the Spark UI. Similar record counts with a much longer duration means the machine is the problem, so speculation helps. An order-of-magnitude larger record count means the partition is the problem, and only repartitioning, salting or adaptive skew handling will fix it.
- Is speculation safe for a job that writes files through Spark?Yes for Spark's own write path: the commit protocol lets only the winning attempt commit its output and discards the loser's staged files, so no duplicates appear. It is not safe for arbitrary external side effects in user code — an HTTP call or a non-transactional append inside foreachPartition will genuinely execute twice.
- Why does speculation do nothing for a stage with only four tasks?The check arms only after spark.speculation.quantile of the stage's tasks have completed, 75% by default, and it needs completed tasks to compute a median duration. With four tasks there is barely a baseline, and by the time three have finished the stage is nearly done anyway, so a duplicate arrives too late to matter.
It is like sending a second courier down a different road when the first is unusually late: helpful if that road is blocked, pointless if the parcel itself is simply enormous.
saying these in an interview costs you the question
- Recommends speculation as the fix for data skew
- Thinks speculation is enabled by default in Spark
- Believes both attempts' outputs get written to the result
- Says speculation fires as soon as any task looks slow
- Assumes it is free because the loser is killed