In spark-submit, what is the difference between --deploy-mode client and --deploy-mode cluster?
answer
- only one process actually moves
- who owns the SparkSession, and where it sits
- kill the terminal and see what happens
- ApplicationMaster container versus your laptop
- shells cannot use one of the two
basics
~20 s--deploy-mode decides where the Spark driver runs. In client mode the driver is the spark-submit process on the submitting machine; in cluster mode the driver runs inside the cluster, in a YARN ApplicationMaster container or a Kubernetes driver pod.
solid answer
~50 sBoth modes run identical executors on the cluster; the flag only moves the **driver**. In `client` mode (the default) the driver lives inside the `spark-submit` JVM on the machine you launched from: driver logs stream to your terminal, the Spark UI is on `localhost:4040`, `collect()` pulls data onto that host's heap, and if the terminal or the CI runner dies the application dies with it. Executors dial back to the driver, so the submitting host must be routable from every cluster node. In `cluster` mode `spark-submit` uploads the artifact and asks the cluster manager to start the driver on the cluster — on YARN that is the ApplicationMaster container, on Kubernetes a driver pod. The launcher can then exit; driver logs become container logs; the driver runs on a sized, monitored, restartable container. Use client for shells and notebooks, cluster for scheduled production jobs.
code
bash · 7 lines# -- driver runs here, in this JVM, on this host
spark-submit --master yarn --deploy-mode client \
--driver-memory 4g --executor-memory 8g app.jar
# -- driver runs in the ApplicationMaster container on the cluster
spark-submit --master yarn --deploy-mode cluster \
--driver-memory 4g --executor-memory 8g app.jargo deeper
Recall the one-line rule: client keeps the driver in your spark-submit process, cluster starts it on the cluster. Know that the shells only work in client mode and that client is the default.
Explain the consequences that follow from driver placement: executors connecting back to spark.driver.host, log destinations, collect() landing on the submitting host, and why --driver-memory must be a flag in client mode.
Show you have been burned: an orchestrator pod OOMed by a client-mode driver, a job killed by a routine gateway restart, an unreachable driver behind NAT. Argue for cluster mode in scheduled pipelines and know how to retrieve driver logs afterwards.
Own the platform default. Decide which submission paths are allowed from where, how driver containers are sized and retried, and how you keep interactive client-mode use off the machines that production scheduling depends on.
## The flag moves one process A Spark application is one driver plus N executors. The driver holds the `SparkSession`, builds the logical and physical plan, owns the DAG scheduler, tracks shuffle map output, and receives the results of `collect()`, `take()` and `count()`. Executors are always started by the cluster manager on cluster nodes. `--deploy-mode` decides only where the driver process runs, and every practical difference between the two modes follows from that one placement decision. ## Client mode Client mode is the default. `spark-submit` parses your arguments and then starts your `main` method *inside its own JVM*. That JVM is the driver. It asks the cluster manager for executors, and each executor opens a connection back to `spark.driver.host` / `spark.driver.port` to receive tasks and return results. Consequences: - **Reachability.** Executors must be able to reach the submitting machine. A laptop behind NAT or a VPN can start a job and then hang, because executors register but cannot connect back. On Kubernetes, client mode needs a headless Service plus an explicit `spark.driver.host` so pods can find the driver. - **Lifetime.** Kill the terminal, lose the SSH session, or let the CI container be reaped, and the driver dies; the application fails immediately. - **Logs and UI.** Driver stdout/stderr go to your terminal; the Spark UI is served from the submitting host on port 4040. - **Result gravity.** `collect()` materialises data in the driver's heap — on a shared edge node, one careless collect can take out everyone's gateway. - **Driver sizing.** Because the driver JVM has already started before your code runs, `spark.driver.memory` set in a `SparkConf` is ignored in client mode; you must pass `--driver-memory`. - **Interactive work.** `spark-shell`, `pyspark`, notebooks and the Thrift server *require* client mode: a REPL needs a driver attached to local stdin, so cluster deploy mode is not applicable to shells. On YARN, client mode still creates an ApplicationMaster, but a stripped-down one whose only job is to negotiate executor containers on behalf of the remote driver. ## Cluster mode In cluster mode, `spark-submit` becomes a launcher only. It stages the application jar or script and any `--jars`/`--files`/`--py-files`, then asks the cluster manager to start the driver on the cluster: - **YARN:** the driver runs *inside* the ApplicationMaster container, sized by `--driver-memory` plus `spark.driver.memoryOverhead`, subject to queue limits. - **Kubernetes:** `spark-submit` creates a driver pod from `spark.kubernetes.container.image`, and that pod then creates the executor pods. - **Standalone:** the driver is launched on one of the worker nodes. Consequences: - **The launcher is disposable.** The job survives the death of the submitting process. On YARN, `spark.yarn.submit.waitAppCompletion` (default `true`) controls whether `spark-submit` blocks reporting status or returns as soon as the application is accepted; Kubernetes has the equivalent `spark.kubernetes.submission.waitAppCompletion`. - **Locality and bandwidth.** The driver sits next to the executors, so scheduling chatter and result traffic stay inside the cluster network. - **Resilience.** On YARN the ApplicationMaster — and therefore the driver — can be retried according to `spark.yarn.maxAppAttempts` (bounded by the cluster's `yarn.resourcemanager.am.max-attempts`). Client mode gives you no such retry: the driver is on your machine. - **Debugging cost.** Nothing useful prints in your terminal beyond a final status. The stack trace lives in the ApplicationMaster's container log, or in `kubectl logs` for the driver pod. Teams meeting this for the first time usually conclude the job "failed silently". - **Local files stop working.** A driver-side path like `/home/me/lookup.csv` read with plain filesystem APIs exists on your gateway, not in the container. Ship it with `--files` or read it from shared storage. ## Choosing Use client mode for interactive exploration, ad-hoc debugging, and short jobs launched by a human who wants to watch the output. Use cluster mode for anything scheduled: an Airflow task, a CI-triggered batch, a long streaming job. The scheduler's worker should not be the machine holding the driver of a six-hour job, both because it is not sized for driver heap and because a routine restart of the scheduler would otherwise kill production jobs. A common hybrid trap: a workflow orchestrator submits in client mode, so the driver lives in a small orchestrator pod. It works for months, then a job with a big broadcast or a wide `collect()` OOMs the orchestrator itself and takes unrelated tasks with it. Moving that job to cluster mode isolates the blast radius to its own container. ## What does not change Executor placement, task scheduling, shuffle behaviour, and every `spark.sql.*` setting are identical in both modes. If a job is slow in one and slow in the other, deploy mode is not the cause — the flag buys you placement, isolation and log location, not performance.
- Why can spark-shell and pyspark not be started with --deploy-mode cluster?A REPL needs the driver attached to the local terminal's stdin and stdout so it can read expressions and print results. In cluster mode the driver runs in a container elsewhere with no interactive channel back, so Spark rejects cluster deploy mode for the shells. Notebook kernels and the Thrift server are client-mode processes for the same reason.
- In client mode on YARN, what does the ApplicationMaster actually do?It is a lightweight executor launcher, not the driver. It registers the application with the ResourceManager and negotiates executor containers on behalf of the remote driver, which keeps running in your spark-submit JVM. In cluster mode the same ApplicationMaster container hosts the driver itself, which is why its memory is sized with --driver-memory.
- Why is spark.driver.memory set in application code ignored in client mode?In client mode the driver JVM is the spark-submit process, and it has already been started with its heap fixed before any of your code executes. Changing the property from a SparkConf therefore cannot resize a running JVM. Pass --driver-memory or set it in the properties file instead; in cluster mode the code-side setting still arrives before the driver container is launched.
- Does deploy mode change how executors are scheduled or how shuffles run?No. Executors are started by the cluster manager and run tasks identically in both modes; shuffle, caching, AQE and every spark.sql setting behave the same. Deploy mode changes driver placement, and therefore network path, log location, failure blast radius and driver-side memory limits — not the distributed execution itself.
Client mode is directing a building site by phone from your office: hang up and work stops. Cluster mode is sending the foreman to live on site — you get a status report instead of a running commentary.
saying these in an interview costs you the question
- Thinks cluster mode moves executors, not the driver
- Says client mode does not use an ApplicationMaster at all
- Expects driver logs in the terminal when running cluster mode
- Claims spark-shell works fine in cluster mode
- Reads a local gateway file path from a cluster-mode driver