skip to content

For a new shared Spark platform, how would you choose between YARN and Kubernetes as the cluster manager?

level: principalimportance: should knowfreq 32%

answer

  1. the engine is identical, the operating model is not
  2. who supplies Python and the native libraries
  3. queues versus namespaces
  4. one of them has no shuffle service
  5. pick what your team can debug at 3am

basics

~20 s

Choose on the surrounding estate, not on Spark. YARN wins where HDFS, Kerberos and queue-based multi-tenancy already exist; Kubernetes wins where dependencies ship as per-job container images and isolation comes from namespaces, at the cost of replacing the external shuffle service.

solid answer

~50 s

Spark runs the same execution engine under both, so decide on the operating model. **YARN** gives you queues with the Capacity or Fair scheduler, node labels, Kerberos delegation tokens, log aggregation, and an external shuffle service that ships as a NodeManager auxiliary service — and it assumes cluster-wide JVM and Python installations that every job shares. **Kubernetes** (GA for Spark since 3.1) makes each job a driver pod plus executor pods built from `spark.kubernetes.container.image`, so dependency versions become per-job rather than platform-wide; isolation is namespaces, ResourceQuota and RBAC; and there is no shuffle service, so dynamic allocation relies on `spark.dynamicAllocation.shuffleTracking.enabled` or a remote shuffle service. Weigh data locality with colocated HDFS DataNodes, your team's existing operational skill, and where the rest of the company's workloads run. Standalone stays a niche for a small dedicated cluster; Mesos has been deprecated since Spark 3.2 and is not a choice for new platforms.

code

bash · 10 lines
bash
# -- YARN: address comes from HADOOP_CONF_DIR, deps from the cluster's install
spark-submit --master yarn --deploy-mode cluster \
  --queue analytics --archives venv.tar.gz#env rollup.jar

# -- Kubernetes: address is the API server, deps come from the image
spark-submit --master k8s://https://api-server:6443 --deploy-mode cluster \
  --conf spark.kubernetes.namespace=analytics \
  --conf spark.kubernetes.container.image=registry/spark-rollup:1.4.0 \
  --conf spark.kubernetes.authenticate.driver.serviceAccountName=spark \
  local:///opt/app/rollup.jar

go deeper

for a junior

Recall that Spark supports several cluster managers — YARN, Kubernetes, standalone and local — and that --master selects which one. The engine and your application code are the same on all of them.

for a middle

Explain concretely what changes between them: driver in an ApplicationMaster container versus a driver pod, cluster-installed Python versus a container image, queues versus namespaces and quotas.

for a senior

Show you have operated at least one of them: how dependencies are shipped, how logs are retrieved, how dynamic allocation is made safe without an external shuffle service, and what breaks first after a migration.

for a principal

Own the decision and its consequences: total operating cost, security and identity model, tenancy fairness, shuffle strategy, and a migration sequenced by workload class rather than a big-bang cutover.

## What a cluster manager actually supplies Spark itself does not schedule machines. It asks a cluster manager for containers, receives them, launches a driver and executors inside them, and gets out of the way. So the choice is not about Spark's performance — Catalyst, the shuffle and the execution model are identical — but about five surrounding concerns: dependency delivery, multi-tenancy and quota, security and identity, shuffle durability, and who operates the thing. ## YARN YARN is the incumbent wherever Hadoop already is. - **Multi-tenancy** is mature: hierarchical queues under the Capacity or Fair scheduler, per-queue capacity and limits, node labels to steer workloads to particular hardware, preemption policies. Twenty teams sharing one cluster is the problem YARN was built for. - **Data locality** is real when NodeManagers sit on the same machines as HDFS DataNodes; Spark can schedule tasks near their blocks. - **Security** is Kerberos: `spark-submit` obtains delegation tokens for HDFS, Hive and HBase and renews them for long-running applications. - **Shuffle** has a first-class answer — the external shuffle service as a NodeManager auxiliary service — which makes dynamic allocation and executor loss cheap. - **Logs** are aggregated by the platform and retrievable with `yarn logs -applicationId`. The cost is the dependency model. Executors run whatever JVM and Python the node has, so upgrading a Python version or a native library is a cluster-wide change negotiated across every team. Fat jars, `--archives` with a packed virtualenv, and careful `--jars` hygiene are the workarounds, and they get old. ## Kubernetes With `--master k8s://https://api-server:6443`, `spark-submit` creates a driver pod from your image and the driver creates executor pods. - **Dependencies become per-job.** The container image pins the JVM, Python, native libraries and connectors. Two teams can run incompatible stacks on the same nodes on the same day. For platforms serving many teams, this is usually the deciding advantage. - **Isolation and quota** come from namespaces, ResourceQuota, LimitRange and RBAC service accounts rather than queues. It is workable but flatter than YARN's hierarchy; scheduling fairness across many pending Spark applications often needs an additional batch scheduler, and "first job grabs everything" is a real failure mode to design against. - **Elasticity** is better when the underlying nodes autoscale, and pods make spot/preemptible capacity natural — provided you enable graceful decommissioning so evictions do not force stage recomputes. - **Shuffle is the gap.** There is no NodeManager to host an auxiliary service, so dynamic allocation depends on `spark.dynamicAllocation.shuffleTracking.enabled` (weaker scale-down) or on adopting a remote/disaggregated shuffle service. Executor-local disk also needs deliberate provisioning rather than being assumed. - **Locality** usually disappears: storage is an object store reached over the network, which changes tuning priorities (larger files, better predicate pushdown) more than it changes the deploy decision. - **Debugging changes shape**: `kubectl logs` on the driver pod, `kubectl describe` for image-pull and scheduling failures, and executor pods that vanish on termination unless you keep them. ## Standalone and the rest Spark's standalone manager is the lightest option: a master and workers, `spark://host:7077`, and per-application core allocation via `--total-executor-cores`. It is a fine choice for a single dedicated cluster owned by one team, and a poor one for multi-tenant fairness or security. `local[*]` remains the right "cluster manager" for tests. Mesos was deprecated in Spark 3.2 and should not anchor a new platform. ## How to actually decide 1. **Where does the rest of the company run?** If every other service is on Kubernetes with established CI, image registries, RBAC and observability, running Spark there means one operational model instead of two. That argument usually beats every technical nuance below it. 2. **How much Hadoop is already load-bearing?** HDFS with Kerberos, Hive metastore integration, existing queue policies and years of tuning are expensive to move; keeping YARN for those workloads while new jobs go to Kubernetes is a legitimate two-platform interim. 3. **How heterogeneous are the dependencies?** Many teams with conflicting Python and library versions push hard toward images. 4. **How shuffle-heavy and how elastic are the workloads?** Shuffle-heavy jobs plus aggressive scale-down plus Kubernetes means you owe a shuffle answer up front, not later. 5. **Who is on call?** Pick the platform your team can debug at 3am. ## Migration reality The migration is not a flag change. `--master` and deploy mode are the easy part; the work is building base images, moving from delegation tokens to service accounts and cloud IAM, replacing queue policies with quotas and a batch scheduler, re-establishing log and history retention, and re-tuning for object storage instead of local HDFS. Plan it per workload class, run both for a period, and keep the same submission wrapper so job authors see one interface.

  • What does Spark Connect change about how clients reach a cluster, whichever manager you pick?
    Spark Connect (Spark 3.4+) splits the client from the driver: a thin client sends unresolved logical plans over gRPC to a Connect server that owns the `SparkSession` on the cluster and returns results as Arrow batches. That removes the fat client-mode driver from notebooks and applications, and it is orthogonal to whether the server runs on YARN or Kubernetes.
  • Which YARN capability has no drop-in replacement on Kubernetes, and how do teams cope?
    The external shuffle service, which on YARN is a NodeManager auxiliary service. On Kubernetes teams either accept weaker scale-down via `spark.dynamicAllocation.shuffleTracking.enabled`, or adopt a remote shuffle service that stores blocks off the executor. Hierarchical queue fairness is the second gap, usually filled with an additional batch scheduler.
  • Would you run a two-platform period, or migrate wholesale?
    Two platforms, deliberately time-boxed. Move new and stateless batch workloads to Kubernetes first, keep Kerberos-and-HDFS-bound jobs on YARN until their storage moves, and hide both behind one submission wrapper so job authors do not learn two interfaces. Wholesale cutovers stall on the handful of jobs nobody understands.

saying these in an interview costs you the question

  • Claims Kubernetes makes Spark jobs inherently faster
  • Assumes an external shuffle service exists on Kubernetes
  • Ignores Kerberos and delegation tokens when leaving YARN
  • Treats the choice as a --master flag change
  • Recommends Mesos for a new Spark platform

context