skip to content

YARN Resource Management

YARN is the cluster's resource broker — the ResourceManager hands out containers, each application gets an ApplicationMaster, and the scheduler decides who waits. Interviewers ask about it because 'my job is stuck in ACCEPTED' is the practical form of this question.

questions

6

In YARN, what do the ResourceManager, NodeManager and ApplicationMaster each do?

level: juniorimportance: must knowfreq 58%

answer

  1. three daemons, two of them cluster-wide
  2. one coordinator per job, not per cluster
  3. who allocates versus who executes
  4. the coordinator itself lives in a container
  5. the scheduler never watches a task

basics

~20 s

YARN splits cluster management three ways: the ResourceManager schedules containers across the whole cluster, a NodeManager runs and monitors containers on each machine, and every application gets its own ApplicationMaster that requests containers and drives that job.

solid answer

~40 s

YARN separates *cluster resource management* from *application logic*. The **ResourceManager** is the single cluster-wide daemon; its Scheduler decides which application gets which container out of the shared pool, and its ApplicationsManager accepts submissions and launches each application's first container. The **NodeManager** runs on every worker: it advertises that machine's memory and vcores, heartbeats to the ResourceManager, launches containers on request, monitors their resource usage, and kills ones that exceed their limits. The **ApplicationMaster** is per-application, not per-cluster — it runs inside an ordinary container, negotiates further containers from the ResourceManager through `allocate()` heartbeats, asks the relevant NodeManagers to start them, tracks progress, and re-requests containers after failures. The ResourceManager never watches individual tasks; that is entirely the ApplicationMaster's job.

code

bash · 4 lines
bash
yarn application -list -appStates RUNNING
yarn node -list -all
yarn application -status application_1699034221_0007
yarn logs -applicationId application_1699034221_0007

go deeper

for a junior

Be able to name the three components and say in one sentence what each does. Interviewers commonly open a Hadoop conversation here, and a confident allocate-versus-execute-versus-coordinate answer is enough.

for a middle

Explain the split as a design decision: why per-application coordinators replaced the single Hadoop 1 JobTracker, and what a container actually is — a memory and vcore lease plus a command, not a Docker container.

for a senior

Show you have operated this: what you look at when an application misbehaves, how the ApplicationMaster's own resource footprint counts against a queue, and what NodeManager auxiliary services such as the external shuffle service buy you.

for a principal

Own the argument for why resource management is separated from application logic at all, and be able to compare that separation with how a container orchestrator divides the same responsibilities when deciding what a platform should run on.

## What problem YARN solves Before YARN, Hadoop's JobTracker did two unrelated jobs at once: it managed cluster resources *and* it tracked every map and reduce task of every job. That made it a scalability bottleneck and locked the cluster into one processing framework. YARN (Yet Another Resource Negotiator, introduced in Hadoop 2 and the model in Hadoop 3.x) splits those responsibilities into a cluster-wide resource broker plus a per-application coordinator, which is why Spark, Tez, Flink and MapReduce can all share one YARN cluster. ## The unit of allocation: a container A YARN **container** is a lease on a slice of one machine's resources — a memory amount and a number of vcores — plus a command line to run inside it. It is not a Docker container by default (though the LinuxContainerExecutor can be configured to use Docker). Everything a YARN application runs, including its own coordinator, runs inside a container. ## ResourceManager The ResourceManager is a single logical daemon (usually deployed as an HA pair with ZooKeeper-based failover). It has two internal pieces: - **Scheduler** — pure allocation. It matches container requests against free capacity according to a pluggable policy (`yarn.resourcemanager.scheduler.class`, Capacity Scheduler by default in Apache Hadoop 3). It does *no* monitoring, *no* task restarts, and offers no guarantees about what runs inside a container. - **ApplicationsManager** — accepts application submissions, validates them, and allocates and launches the first container of each application, the one that hosts the ApplicationMaster. It also restarts a failed ApplicationMaster within the configured attempt limit. The ResourceManager knows about applications and containers. It does not know that a Spark job has 4,000 tasks or that a MapReduce job is in its reduce phase. ## NodeManager One NodeManager per worker machine. Its responsibilities: - Advertise the machine's capacity — `yarn.nodemanager.resource.memory-mb` and `yarn.nodemanager.resource.cpu-vcores` — and heartbeat liveness and container status to the ResourceManager. - **Localize** resources: download the job's jars, archives and files from HDFS to the local disk before starting a container. - Launch containers through a ContainerExecutor, and monitor their memory and CPU use, killing a container that exceeds the memory it was granted. - Run **auxiliary services** that outlive individual containers — the MapReduce shuffle handler, and Spark's external shuffle service when it is enabled. - Aggregate container logs to HDFS when `yarn.log-aggregation-enable` is on, which is what makes `yarn logs -applicationId <id>` work after an application finishes. A NodeManager never decides *which* application gets resources. It only executes and polices what the ResourceManager granted. ## ApplicationMaster This is the piece people miss. Every submitted application gets its own ApplicationMaster, running in a container on some NodeManager chosen by the ResourceManager — never on the ResourceManager host, and (in cluster deploy mode) never on the client machine. It is framework-specific code: MapReduce ships `MRAppMaster`, Spark ships its own, and in Spark's cluster deploy mode the driver itself runs inside the ApplicationMaster container. The ApplicationMaster: 1. Registers with the ResourceManager. 2. Sends periodic `allocate()` heartbeats carrying resource requests (how many containers, how much memory and how many vcores each, and locality preferences such as "on this host" or "on this rack"). 3. Receives allocated containers in the heartbeat responses and contacts the owning NodeManagers to start processes in them. 4. Monitors its own tasks, re-requesting containers for the failed ones and applying its own retry and speculation policy. 5. Unregisters and exits when the application finishes, releasing its containers. Because the retry logic lives here, a task failure never touches the ResourceManager. Scaling the cluster scales the number of ApplicationMasters rather than the load on one central tracker. ## End-to-end submission A client submits an application to the ResourceManager. The ResourceManager picks a node with free capacity and asks that NodeManager to launch the ApplicationMaster container. The ApplicationMaster starts, registers, and begins asking for containers. As containers are granted, the ApplicationMaster tells the relevant NodeManagers to launch the actual work. The application passes through the states NEW, SUBMITTED, ACCEPTED (accepted by the scheduler, waiting for its ApplicationMaster container), RUNNING, and finally FINISHED, FAILED or KILLED. ## Vocabulary traps A YARN **container** is a resource lease, not a Docker image. A MapReduce **task** is a whole mapper or reducer, which occupies one container; a Spark **task** is one partition's work inside an executor, and one Spark executor occupies one YARN container and runs many tasks. And a YARN application's coordinator is the ApplicationMaster — a Spark driver and a YARN ApplicationMaster coincide only in cluster deploy mode.

  • Where does the ApplicationMaster actually run, and why does that matter?
    In an ordinary container on some NodeManager, chosen by the ResourceManager — it is the application's first container. It matters because the ApplicationMaster consumes cluster resources like any other container, counts against the queue's ApplicationMaster budget, and dies with its node: if that machine is lost, the whole application attempt is lost and must be retried.
  • What is the ResourceManager deliberately not responsible for?
    Task-level monitoring, retries, speculation and progress reporting. Its Scheduler only matches container requests to free capacity; it makes no guarantee about what runs inside a container or whether it succeeds. Pushing that work into per-application ApplicationMasters is precisely what removed the Hadoop 1 JobTracker bottleneck.
  • What are NodeManager auxiliary services and why do they exist?
    Long-lived services hosted in the NodeManager process, outside any container, so their data survives the container that produced it. The MapReduce shuffle handler serves map output to reducers after the mapper's container exits, and Spark's external shuffle service does the same for executors — which is what allows executors to be removed by dynamic allocation without losing shuffle files.

The ResourceManager is a building's letting agent who only decides who gets which rooms; the NodeManager is the caretaker of one floor who unlocks doors and throws out tenants who overrun their space; the ApplicationMaster is the project lead each tenant sends in to actually run their work in the rooms they were given.

saying these in an interview costs you the question

  • Says the ResourceManager launches and monitors every individual task
  • Thinks one ApplicationMaster coordinates the whole cluster
  • Claims the ApplicationMaster runs on the ResourceManager host
  • Says the NodeManager decides which application gets resources
  • Confuses YARN with HDFS and thinks YARN stores the data

context

open as a page

Why does a YARN NodeManager kill a container for exceeding physical memory limits?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Each YARN container is granted a fixed memory amount, and the NodeManager monitors the container's whole process tree against it. When total resident memory crosses the grant the container is killed, because YARN protects the node from oversubscription rather than letting the machine swap or OOM.

open as a page

How do YARN's Capacity Scheduler and Fair Scheduler differ when sharing one cluster?

level: middleimportance: should knowfreq 45%

basics

~20 s

Both divide a YARN cluster into hierarchical queues, but the Capacity Scheduler starts from guaranteed percentages per queue with an elastic ceiling, while the Fair Scheduler starts from equal sharing among running applications and pulls resources back toward each queue's fair share over time.

open as a page

A YARN application sits in ACCEPTED state for 20 minutes and never runs — how do you diagnose it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

ACCEPTED means YARN took the application but has not launched its ApplicationMaster container yet. The cause is always missing capacity: the queue is full, its ApplicationMaster budget is exhausted, a per-user limit binds, or no healthy node can host the requested container size.

open as a page

How would you design YARN queues for a cluster shared by teams with different SLAs?

level: principalimportance: should knowfreq 32%

basics

~20 s

Shape queues around workload classes and SLAs rather than around org charts: give SLA-bound pipelines guaranteed capacity with limited borrowing, let ad-hoc and batch work share a large elastic queue, cap concurrency and per-user limits, and enable preemption only where redoing work is cheap.

open as a page

In YARN, what happens when an ApplicationMaster container fails, and what caps the retries?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The ResourceManager notices the ApplicationMaster is gone and starts a fresh attempt in a new container elsewhere, up to yarn.resourcemanager.am-max-attempts (default 2). When attempts run out the whole application is marked FAILED, and by default its earlier containers are killed too.

open as a page