When is YARN still the right cluster manager for Spark or Flink instead of Kubernetes?
answer
- who already owns the cluster?
- where does the data actually live?
- locality only counts if storage is local
- no container image to build
- Kerberos and queues already configured
basics
~20 sYARN still wins on an existing Hadoop estate: data already in HDFS on the same machines, Kerberos and queue policies already configured, and no container images to build or maintain. Greenfield platforms reading object storage normally choose Kubernetes.
solid answer
~50 sYARN is Hadoop's cluster resource manager, and both Spark (`--master yarn`) and Flink still ship it as a deployment target alongside Kubernetes. I would keep it when the estate is already Hadoop: the data sits in HDFS on the same worker machines so locality-aware placement is real, Kerberos authentication and the queue policies that arbitrate between teams already exist, and the operations team knows how to run it. Nothing has to be containerised — the engine ships as a JAR into a container process the NodeManager launches from the node's own environment. I would move to Kubernetes when the data has left HDFS for object storage, so locality buys nothing; when the same platform must also run services, not just batch; or when immutable images and per-job dependency isolation matter more than the Hadoop integration. The migration is not a flag change: image builds, secret handling, and shuffle/dynamic-allocation behaviour all differ.
code
bash · 6 lines# YARN: the engine ships as a JAR into a NodeManager-launched process
spark-submit --master yarn --deploy-mode cluster app.py
# Kubernetes: the same job now needs a published container image
spark-submit --master k8s://https://api-server:6443 --deploy-mode cluster \
--conf spark.kubernetes.container.image=registry.example.com/spark:3.5.1 app.pygo deeper
Know that YARN is Hadoop's cluster resource manager and that Spark and Flink can be submitted to it as one deployment target among several, Kubernetes being the other common one. You are not expected to choose between them.
Be ready to explain why a Hadoop estate favours YARN — engine processes launched next to the HDFS data they read, no container image required — and why moving the data to object storage removes that advantage.
Show you have made this call on a real platform: state the decision rule, then name the concrete migration work — image builds, identity and secrets, shuffle and dynamic-allocation behaviour, log access — rather than presenting it as a configuration change.
Own the strategy: whether to run two schedulers or consolidate, what stranded capacity and a duplicated on-call rotation cost, how team skills and vendor support direction weigh against a multi-quarter migration, and how you would sequence it so no workload is stranded mid-flight.
## What this question is really asking YARN (Yet Another Resource Negotiator) is the cluster resource manager of the Hadoop 3.x stack: a central ResourceManager grants resource containers, a NodeManager on each machine launches and supervises the processes inside them, and per-application coordination runs in an ApplicationMaster. Spark and Flink do not depend on YARN — they use it as one of several *deployment targets*, next to Kubernetes and standalone clusters. So the interview question is not "how does YARN work internally"; it is a platform-choice question: given that Kubernetes has become the default substrate for new data platforms, when does an experienced engineer still say "YARN"? A weak answer treats YARN as obsolete Hadoop baggage. A strong answer names the specific properties that still favour it, and — just as important — names the conditions under which those properties have already evaporated. ## What still argues for YARN **Co-located storage.** The classic Hadoop deployment puts HDFS DataNodes and YARN NodeManagers on the same physical machines, so a task can be scheduled where the input blocks already live and read them from local disk. If your tables genuinely live in HDFS on those nodes, that locality is a real saving on a shuffle-heavy or scan-heavy workload. **Existing security and tenancy configuration.** A mature Hadoop estate has Kerberos identities, delegation tokens, HDFS permissions, and queue policies that decide which team's jobs get resources when the cluster is full. Reproducing an equivalent policy set on a fresh platform is months of work that delivers no new capability. **No container image supply chain.** On YARN the engine is a JAR (plus whatever the node's environment provides). Nobody has to build, scan, sign, store, and garbage-collect an image per engine version. For a small team, that is real operational relief. **Operational familiarity.** The people on call already know how to read that cluster. Migration replaces a well-understood failure surface with an unfamiliar one. ## What argues against it **The data usually moved.** Once tables live in cloud object storage, every read is remote regardless of where the executor lands, and locality-aware placement has nothing local to match. That removes YARN's headline technical advantage — and it is the single most common reason the answer flips. **One platform instead of two.** Most organisations already run Kubernetes for their services. Running a second scheduler purely for data jobs means two capacity pools, two on-call rotations, two upgrade cadences and stranded capacity in each. **Dependency isolation.** On YARN, jobs inherit a shared node environment, and classpath conflicts between the engine's dependencies and the job's are a chronic irritation. An immutable image per job is a genuinely better isolation story. **Ecosystem direction.** Investment — vendor distributions, tooling, operators, examples — has shifted toward Kubernetes. Spark on Kubernetes reached general availability in Spark 3.1 and both engines continue to support it as a first-class target in Spark 3.5/4.x and Flink 1.19/2.x. Neither has removed YARN support, but new work targets Kubernetes. ## Why the migration is not a flag change Engineers who have not done it assume you swap `--master yarn` for `--master k8s://…` and stop. Concretely, what also changes: - **Packaging.** Kubernetes needs a container image; on Spark you must set `spark.kubernetes.container.image`, and something has to build and publish it per engine version. - **Elastic executors.** Spark's dynamic allocation on YARN traditionally relies on an external shuffle service so an executor holding shuffle output can still be released. On Kubernetes the common arrangement instead keeps executors that hold shuffle data alive, via `spark.dynamicAllocation.shuffleTracking.enabled`. The scaling behaviour you observe is therefore different. - **Identity and secrets.** Kerberos tickets and delegation tokens give way to service accounts, mounted secrets and cloud workload identity. - **Tenancy.** Queue policies become namespaces, quotas and priority classes — a different model with different fairness behaviour, not a translation. - **Observability.** Where you fetch driver and executor logs, and what a "container killed for exceeding memory" event looks like, both move. ## How to answer it in an interview State the decision rule first — *where does the data live, and does a Hadoop estate already exist?* — then give the two branches, then show you know the migration has teeth. Avoid the two failure modes: dismissing YARN as MapReduce-only (Spark and Flink run on it perfectly well), and defending it on locality grounds when the data has already moved to object storage.
- Your tables have moved from HDFS to cloud object storage but the jobs still run on YARN. What has that cost you?Most of YARN's technical advantage. Locality-aware placement has nothing local to match, because every read crosses the network to the object store regardless of which machine the task lands on. What remains is inertia value: existing Kerberos setup, existing queue policies, and an operations team that knows the platform. That is a legitimate reason to defer a migration, but it is an organisational argument, not a performance one.
- What changes about Spark's dynamic allocation when a job moves from YARN to Kubernetes?On YARN, dynamic allocation is normally paired with an external shuffle service, so an idle executor can be released while its shuffle output stays served by the node. Kubernetes deployments commonly enable `spark.dynamicAllocation.shuffleTracking.enabled` instead, which keeps executors holding referenced shuffle data alive rather than externalising it. The result is that executors are reclaimed less aggressively, so scale-down behaviour and cost profile differ even with identical job code.
- A team wants to keep YARN because "Kubernetes cannot schedule big data workloads". How do you respond?I would push back on the premise. Both Spark and Flink ship first-class Kubernetes deployment support and it is where new investment goes. The honest arguments for staying are HDFS co-location, existing Kerberos and queue policy, and operational familiarity — not an inability of Kubernetes to run the workload. Framing it accurately matters, because it turns the decision into a cost-and-timing question rather than a technical veto.
saying these in an interview costs you the question
- Says YARN can only run MapReduce, not Spark or Flink
- Claims data locality still helps when tables live in object storage
- Assumes migrating is just swapping the --master flag
- Treats a YARN container and a Kubernetes Pod as the same thing
- Says Kubernetes replaced YARN entirely and both engines dropped support