A YARN application sits in ACCEPTED state for 20 minutes and never runs — how do you diagnose it?
answer
- it was accepted, so nothing crashed
- the first container is the missing one
- ten percent of the queue, by default
- one user can throttle themselves
- a full local disk removes a whole node
basics
~20 sACCEPTED means YARN took the application but has not launched its ApplicationMaster container yet. The cause is always missing capacity: the queue is full, its ApplicationMaster budget is exhausted, a per-user limit binds, or no healthy node can host the requested container size.
solid answer
~40 sACCEPTED means the scheduler accepted the application but cannot yet place its **ApplicationMaster** container, so the answer is always "something is refusing capacity". Work through it on the ResourceManager web UI and with `yarn application -list -appStates ACCEPTED`. Check, in order: is the target queue at its `capacity` or `maximum-capacity`; has the queue hit `yarn.scheduler.capacity.maximum-am-resource-percent` (0.1 by default), so all its headroom is already spent on other ApplicationMasters; is `user-limit-factor` (default 1) capping this user; are there healthy NodeManagers at all — unhealthy nodes from failed disk checks silently remove capacity; and is the requested container larger than any node has free, or larger than `yarn.scheduler.maximum-allocation-mb`. The classic deadlock is many small applications whose ApplicationMasters fill the queue's 10% budget while each waits for worker containers that will never be free.
code
bash · 4 linesyarn application -list -appStates ACCEPTED
yarn queue -status etl
yarn node -list -all # look for UNHEALTHY / LOST
yarn application -kill application_1699034221_0031go deeper
Know that ACCEPTED means YARN has the application but has not started it yet, and that it points at cluster capacity rather than a bug in your code.
Explain that the missing container is the ApplicationMaster's own, and name at least two limits that can withhold it — queue capacity and the per-queue ApplicationMaster percentage.
Walk the diagnosis in order across the scheduler page, node list and queue configuration, and separate a genuine shortage from a self-inflicted one such as abandoned sessions or a burst of concurrent submissions.
Frame it as a capacity and admission-control design question: how many concurrent applications a queue is sized for, who is allowed to burst, and whether preemption or hard concurrency caps is the right instrument.
## What ACCEPTED actually means A YARN application moves NEW → NEW_SAVING → SUBMITTED → ACCEPTED → RUNNING. **ACCEPTED** means the ResourceManager has validated and persisted the submission and handed it to the scheduler, but the scheduler has not yet been able to allocate and launch the application's *first* container — the one that hosts the ApplicationMaster. The application only reaches RUNNING once that ApplicationMaster has started and registered. So a stuck ACCEPTED is never an application bug; it is a scheduling refusal. ## The diagnosis order **1. Which queue did it land in, and is that queue full?** Open the ResourceManager UI's scheduler page (or `yarn queue -status <queue>`) and read Used Capacity against Configured Capacity and Max Capacity. Under the Capacity Scheduler a queue is guaranteed `yarn.scheduler.capacity.root.<q>.capacity` and may borrow up to `maximum-capacity` only when other queues are idle. If neighbours are busy, your queue is pinned at its guarantee even though the cluster looks partly free. Also confirm the queue is the one you intended — queue mappings (`yarn.scheduler.capacity.queue-mappings`) can route a user somewhere unexpected. **2. Is the queue's ApplicationMaster budget exhausted?** This is the single most common cause and the least obvious. `yarn.scheduler.capacity.maximum-am-resource-percent` defaults to **0.1**: at most 10% of a queue's resources may be spent on ApplicationMaster containers. That cap exists precisely to prevent a queue filling with coordinators and no workers. When it binds, the scheduler will not start another ApplicationMaster no matter how much worker capacity is free, and applications pile up in ACCEPTED. The scheduler page shows "Max Application Master Resources" and "Used Application Master Resources" per queue — compare them. **3. Is a per-user limit binding?** `yarn.scheduler.capacity.root.<q>.user-limit-factor` defaults to **1**, meaning a single user cannot consume more than the queue's *configured* capacity even when the queue is allowed to elastically borrow beyond it. One user submitting a burst of applications throttles themselves while the queue's max capacity sits unused. `minimum-user-limit-percent` similarly divides a queue among active users. **4. Is there any capacity at all?** `yarn node -list -all` shows NodeManagers and their states. Nodes go **UNHEALTHY** when the disk health checker trips — by default a node is marked unhealthy when a local directory exceeds `yarn.nodemanager.disk-health-checker.max-disk-utilization-per-disk-percentage` (90%) — and an unhealthy node contributes zero capacity while still appearing in the cluster. A cluster that filled its local disks with un-aggregated logs can lose most of its capacity this way. Also check that NodeManagers are actually connected: a network partition or a mass restart leaves the ResourceManager with few or no nodes. **5. Is the request unsatisfiable by shape?** The scheduler rounds every request up to a multiple of `yarn.scheduler.minimum-allocation-mb` (default 1024) and rejects anything above `yarn.scheduler.maximum-allocation-mb` (default 8192) — an oversized request usually fails at submit with an InvalidResourceRequestException rather than hanging. The subtler version is a request that is legal but larger than any single node currently has free: a 60 GB ApplicationMaster on nodes advertising 64 GB that are each 20% used will wait indefinitely. Node labels are the same trap — an application pinned to a labelled partition with no healthy nodes in it waits forever. ## The classic deadlock A scheduled workflow fires fifty small applications at once into a queue. Each gets its ApplicationMaster started until the 10% ApplicationMaster budget is spent; those ApplicationMasters then request worker containers, but the remaining 90% is consumed by *other* applications' workers that are themselves progressing slowly. New submissions sit in ACCEPTED behind the ApplicationMaster cap. Nothing is broken; the cluster is simply oversubscribed at the coordinator level. Raising `maximum-am-resource-percent` blindly makes it worse — you get more coordinators competing for the same workers. The real fixes are throttling concurrent submissions, giving bursty pipelines their own queue, or increasing capacity. ## Remediation Short term: kill abandoned applications (`yarn application -kill <id>`) — long-idle Spark shells and notebook sessions holding containers are a frequent culprit. Medium term: separate workloads into queues so one team's burst cannot occupy another's guarantee, enable preemption so a queue can reclaim its guaranteed share, and fix unhealthy nodes. Long term: cap concurrency per queue with `yarn.scheduler.capacity.root.<q>.maximum-applications` and size the ApplicationMaster budget for the number of concurrent applications you actually expect. ## What it is not ACCEPTED is not "the job is running slowly", and it is not a data problem. If the application reaches RUNNING and then stalls, you have left YARN's territory and are debugging the framework — a slow stage, a skewed task, a wedged coordinator — not the resource broker.
- Why does raising maximum-am-resource-percent often fail to fix the pile-up?Because the cap is not the shortage, it is the guard against one. Admitting more ApplicationMasters just adds more coordinators competing for the same scarce worker containers, and can deadlock a queue where every resource is held by a coordinator waiting for workers. Fix the concurrency or the capacity instead, and only raise the percentage when you genuinely run many small applications with tiny worker footprints.
- How would you tell a queue-capacity problem apart from an unhealthy-node problem?Read the scheduler page and the node list together. If Used Capacity is at Max Capacity, it is a queue limit. If the queue looks idle but the cluster's total advertised memory is far below what the hardware should give, NodeManagers are unhealthy or disconnected — check the node list for UNHEALTHY entries and the NodeManager health report, usually a local disk over the utilisation threshold.
- The application reaches RUNNING but its workers never start. Is that the same problem?It is the same shortage one level down. RUNNING means the ApplicationMaster launched; the coordinator is now requesting worker containers and getting none, so you check the same queue capacity, user limit and node health, but the application's own UI shows the pending request shape — how large a container it wants and with what locality — which is often what makes it unschedulable.
saying these in an interview costs you the question
- Assumes the application code or its data is at fault
- Immediately raises maximum-am-resource-percent without looking at usage
- Thinks ACCEPTED means the job is running and merely slow
- Never checks for unhealthy NodeManagers removing capacity
- Ignores that the queue may be at its guarantee while the cluster looks free