In a Spark executor, what does spark.memory.fraction control and how do execution and storage share it?
answer
- two consumers, one shared budget
- a fixed slice is reserved first
- only one side may evict the other
- 0.6 of the heap after 300 MB
basics
~20 sspark.memory.fraction (default 0.6) sizes the unified pool for execution and storage out of the executor heap left after a fixed 300 MB reservation. Execution and storage borrow from each other on demand; only storage's protected floor is safe from eviction.
solid answer
~50 sAn executor's JVM heap is `spark.executor.memory`. Spark subtracts a fixed 300 MB reservation, then takes `spark.memory.fraction` (default 0.6) of what remains as the **unified pool** shared by execution and storage. Execution memory is scratch space for shuffles, sorts, joins and aggregations; storage memory holds cached blocks and broadcast data. Inside the pool, `spark.memory.storageFraction` (default 0.5) is not a hard partition — it is the share of storage that execution may **not** evict. Either side can borrow the other's idle space. The asymmetry matters: execution can evict borrowed-from-storage blocks down to that floor, but storage can never evict execution memory — it waits or spills. The other 40% of the heap is *user memory*: your UDF objects, closures and Spark's own metadata. Since Spark 3.0 the old static memory manager is gone, so this is the only model.
code
properties · 5 linesspark.executor.memory 10g
spark.memory.fraction 0.6
spark.memory.storageFraction 0.5
# unified pool = (10240m - 300m) * 0.6 ~= 5964m
# eviction-immune storage floor = 5964m * 0.5 ~= 2982mgo deeper
Be able to say that an executor's memory is not one undifferentiated blob: part of it holds cached data, part is scratch space for shuffles and joins, and part belongs to your own code.
You are expected to name spark.memory.fraction and spark.memory.storageFraction, state their defaults, and explain the 300 MB reservation and the asymmetric borrowing rule between execution and storage.
Show that you read the split from real symptoms — heavy spill next to a fully-cached RDD, or an OutOfMemoryError inside a UDF — and that you know which of the two knobs, if either, is the right response versus adding partitions or executors.
Own the position that these fractions redistribute rather than create memory, and set a platform default that leaves user memory intact instead of letting each team tune spark.memory.fraction upward until executors die.
## The heap you are dividing When an executor starts, `spark.executor.memory` becomes the JVM's `-Xmx`. Everything discussed below is a division of *that* heap; container overhead, off-heap buffers and Python workers live outside it and are budgeted separately by `spark.executor.memoryOverhead`. ## Step 1 — the 300 MB reservation Spark reserves a fixed 300 MB of system memory before doing any arithmetic. It is not configurable in production (only a testing property changes it), and it exists so a tiny executor still has room for the JVM's own structures. A consequence people hit in local experiments: an executor heap below roughly 450 MB refuses to start, because Spark demands at least 1.5x the reservation. ## Step 2 — the unified pool Of the heap *minus* 300 MB, `spark.memory.fraction` (default **0.6**) becomes the unified memory pool. On a 10 GB executor that is roughly `(10240 - 300) * 0.6` ≈ 5.96 GB. This pool is shared by two consumers: - **Execution memory** — transient working space: shuffle write and read buffers, sort buffers, hash tables for joins and aggregations, and the Tungsten binary rows those operators build. - **Storage memory** — cached and persisted blocks (`cache()`/`persist()`), plus broadcast variables and the broadcast side of a broadcast hash join. ## Step 3 — storageFraction is a floor, not a wall `spark.memory.storageFraction` (default **0.5**) does **not** cut the pool in half permanently. It defines the fraction of the pool that storage is *immune from eviction* within. The borrowing rules are deliberately asymmetric: - If storage is idle, execution may take the free space — and if execution later needs more, it may **evict** cached blocks, but only down to the `storageFraction` floor. - If execution is idle, storage may take the free space — but storage can **never** evict execution memory. A caching task that finds the pool full simply cannot cache the block, and (depending on the storage level) drops it or writes it to disk. This asymmetry is intentional. Evicting a cached block costs a recomputation or a disk read; evicting an in-flight sort buffer would cost correctness or a forced spill in the middle of an operator. ## Step 4 — user memory The remaining `1 - spark.memory.fraction` (about 40%) of the usable heap is *user memory*. Spark does not account for it at all. It holds objects your code allocates — the `HashMap` you build inside a `mapPartitions`, the list a UDF accumulates, deserialized records, and Spark-internal metadata that is not tracked by the memory manager. Jobs whose UDFs allocate heavily fail here, and no amount of raising `spark.memory.fraction` helps — raising it makes user memory *smaller*. ## Step 5 — how the pool is divided among tasks Execution memory is not per-task-reserved. An executor runs `spark.executor.cores` tasks concurrently, and the execution pool is shared dynamically: a task is guaranteed at least `1/(2N)` of the pool and may acquire at most `1/N`, where N is the number of currently active tasks. Two consequences follow. First, running more cores per executor gives each task a smaller slice, which is a very common cause of spill. Second, if you cache aggressively you shrink what execution can borrow, and the same job that ran clean yesterday spills today. ## What you actually change - **Cached data dominates and execution spills** → lower `spark.memory.storageFraction` so execution may evict more, or cache less, or use `MEMORY_AND_DISK_SER`. - **UDF-heavy job throwing `OutOfMemoryError` in user code** → *lower* `spark.memory.fraction` to leave more user memory, or fix the allocation. - **Everything is tight** → more executors or a larger heap; the fractions only redistribute, they do not create memory. ## What this model is not Before Spark 1.6 the split was static (`spark.storage.memoryFraction` / `spark.shuffle.memoryFraction`), and `spark.memory.useLegacyMode` let you opt back in. **All of that was removed in Spark 3.0.** Quoting those keys in an interview marks the answer as a decade out of date. Likewise, this on-heap model is separate from the off-heap pool, which exists only when `spark.memory.offHeap.enabled` is `true` and has its own fixed size.
- If execution evicts a cached block that was sitting in borrowed storage space, what happens to that data?It depends on the storage level. With MEMORY_AND_DISK the block is written to the executor's local disk and read back later. With MEMORY_ONLY it is simply dropped, and the next action that needs that partition recomputes it from lineage. Either way the job stays correct; it just gets slower, and the Storage tab in the Spark UI shows the cached fraction falling below 100%.
- Would raising spark.memory.fraction to 0.9 be a safe way to stop a job from spilling?Rarely. It shrinks user memory to 10% of the heap, so any job whose UDFs, closures or deserialized records allocate on the heap starts throwing OutOfMemoryError in user code instead of spilling gracefully. Spilling is a controlled slowdown; running out of user memory kills the executor. Tune the number of partitions or executor cores first.
- Where does the broadcast side of a broadcast hash join live in this model?In storage memory on each executor, alongside cached blocks — which is why an over-large broadcast both fails the join and evicts your cache. spark.sql.autoBroadcastJoinThreshold bounds what the planner will broadcast automatically, but it compares against estimated sizes, so an under-estimate can still push a much larger relation into storage memory.
One shared desk with two people: the cache spreads out when the shuffle worker is idle, but the shuffle worker can sweep the cache's papers off the desk down to an agreed reserved corner, while the cache can never sweep the shuffle worker's papers.
saying these in an interview costs you the question
- Says storageFraction hard-partitions the pool in half
- Claims storage can evict execution memory when it needs space
- Quotes spark.storage.memoryFraction, removed in Spark 3.0
- Thinks spark.memory.fraction covers the whole executor container
- Suggests raising memory.fraction to 0.9 with no mention of user memory