In Flink, how do task slots relate to a job's parallelism and to slot sharing?
answer
- a worker offers a fixed number of these
- one slice of the pipeline per slot
- memory is divided, CPU is not
- widest operator sets the requirement
- a named group opts out of sharing
basics
~20 sA Flink task slot is one TaskManager's unit of scheduling, and by default subtasks of different operators in the same job share a slot. So a job needs as many slots as its highest operator parallelism, not the sum of all parallelisms.
solid answer
~50 sEach TaskManager offers `taskmanager.numberOfTaskSlots` slots (default 1); a slot gets an equal share of the TaskManager's managed memory but **no CPU isolation** — slots share the JVM's cores and heap. By default all subtasks of one job belong to the slot sharing group `default`, and Flink will place one subtask of *each* operator into the same slot. A slot therefore holds a full vertical slice of the pipeline — source, map, window, sink — so the number of slots a job needs equals the **maximum parallelism of any single operator**, not the sum across operators. That keeps the arithmetic simple and mixes cheap and expensive operators in one slot for better utilization. Calling `slotSharingGroup("heavy")` on an operator pulls it into its own group so it gets dedicated slots; then the requirement is the sum over groups. If the cluster has fewer slots than required, the job never starts.
code
yaml · 3 linestaskmanager.numberOfTaskSlots: 4
taskmanager.memory.process.size: 8g
parallelism.default: 4go deeper
Know that a TaskManager offers a configured number of slots and that a job's parallelism is how many parallel subtasks each operator runs. Being able to read the Web UI's available-slot count already helps.
Explain the slot-sharing rule and derive the slot requirement from a job graph on the spot: maximum operator parallelism, not the sum. Know that managed memory is split per slot while CPU is not.
Show judgment on when to break sharing with a named slot sharing group, and on TaskManager shape — many thin workers versus few fat ones — in terms of GC behaviour and failure domain.
Own the fleet-level convention: a default slots-per-core ratio, a container shape, and rules for when teams may carve out dedicated slot sharing groups, since every carve-out lowers utilization across the platform.
## What a slot actually is A **task slot** is the unit in which a Flink TaskManager offers resources. A TaskManager advertises a fixed number of them, set by `taskmanager.numberOfTaskSlots` (default `1`). A slot is a *scheduling* unit and a *memory* unit, not a CPU unit: - The TaskManager's **managed memory** (the off-heap pool Flink uses for RocksDB and for batch operators, sized by `taskmanager.memory.managed.fraction`, default 0.4 of Flink memory) is divided equally among the slots. - CPU and JVM heap are **not** partitioned. Subtasks in different slots of the same TaskManager compete for the same cores and share one garbage collector. That is why the usual sizing rule of thumb is to give a TaskManager about as many slots as it has CPU cores, and why one runaway subtask can slow down its slot-neighbours. ## Slot sharing By default every operator of a job belongs to the slot sharing group named `default`, and Flink allows subtasks from *different* tasks of the same job to occupy the same slot. It will never put two subtasks of the *same* task into one slot — that would defeat parallelism. The consequence is that a slot ends up holding one parallel slice of the whole pipeline. For a job `source(4) -> map(4) -> keyBy/window(4) -> sink(4)`, each of the four slots holds one source subtask, one map subtask, one window subtask and one sink subtask. The job needs **4** slots, not 16. More generally: **slots required = the highest parallelism among the job's operators**. If the source runs at 2 (because the Kafka topic has 2 partitions) and the window runs at 8, the job needs 8 slots; the two source subtasks sit in two of them and the other six slots simply have no source subtask. Two benefits follow. First, the required slot count is easy to reason about and equals the job's parallelism. Second, resource-hungry and resource-cheap operators are mixed: a slot rarely contains only expensive work or only idle work, so utilization evens out. ## Breaking sharing on purpose `someStream.slotSharingGroup("heavy")` moves an operator (and everything chained after it, until another group is declared) into its own group. Subtasks from different groups never share a slot, so the cluster requirement becomes the **sum over groups of each group's maximum parallelism**. You reach for this when one operator holds a very large RocksDB state and you want it to own a slot's managed memory outright, or when an operator makes blocking external calls and you do not want it stealing time from a latency-sensitive neighbour. The price is more slots and lower average utilization. ## Chaining: a different fusion Slot sharing is often confused with **operator chaining**. Chaining fuses *consecutive* operators that have a forward (one-to-one) connection and the same parallelism into a single **task** that runs in **one thread**, passing records by method call with no serialization. Slot sharing places subtasks of *different* tasks in the same slot, each still its own thread. Chaining is controlled with `disableChaining()`, `startNewChain()` and `env.disableOperatorChaining()`, and it is what makes the Web UI show `Source: kafka -> map -> filter` as a single box. Disabling it is a debugging aid — you get separate metrics per operator — and it costs serialization and thread handoffs. ## When there are not enough slots Flink does not quietly reduce parallelism. If a job requires more slots than the cluster can supply, the JobMaster requests them, waits, and after `slot.request.timeout` fails the job with a `NoResourceAvailableException`-style error. In native Kubernetes or YARN deployments, Flink's ResourceManager first tries to start more TaskManagers; in a standalone cluster it can only wait. The Web UI's cluster overview showing "Available Task Slots: 0" next to a job stuck in scheduling is the everyday symptom. ## Practical sizing Two TaskManagers with 8 slots each and one with 16 slots offer the same 16 slots, but not the same behaviour: fewer, fatter TaskManagers mean fewer JVMs to manage, more shared heap, larger GC pauses and a bigger failure domain (losing one TaskManager kills more subtasks and forces a restart of the job). Many, thinner TaskManagers isolate better and lose less on failure, at the cost of more JVM overhead and more network connections. A common middle ground is one TaskManager per node with slots equal to the cores allocated to it.
- A job's source runs at parallelism 2 and its window operator at 8. How many slots does the job need?Eight. With default slot sharing, the requirement is the maximum operator parallelism, not the sum. Eight slots each hold one window subtask; two of them also hold a source subtask and the remaining six hold none. The source's parallelism 2 is usually itself a ceiling imposed by the number of Kafka partitions.
- Do two subtasks in the same slot get isolated CPU?No. A slot divides the TaskManager's managed memory equally, but CPU cores and JVM heap are shared across all slots in that TaskManager, and all subtasks share one garbage collector. A subtask making slow blocking calls or triggering long GC pauses will affect its slot-neighbours, which is one reason to size slots roughly to cores and to isolate genuinely hostile operators into their own slot sharing group.
- How does operator chaining differ from slot sharing?Chaining fuses consecutive operators with a forward connection and equal parallelism into one task running in one thread, handing records over by method call with no serialization. Slot sharing places subtasks of different tasks into the same slot, each still on its own thread. Chaining is a per-thread optimization; slot sharing is a placement rule.
A slot is a workbench that gets one worker from each stage of the line, not one bench per stage. You need as many benches as the busiest stage has workers, not as many as all stages combined.
saying these in an interview costs you the question
- Adding up every operator's parallelism to size the cluster
- Claiming slots isolate CPU as well as memory
- Treating operator chaining and slot sharing as the same mechanism
- Expecting Flink to run at lower parallelism when slots are short
- Setting one slot per TaskManager on a multi-core node without reason