skip to content

On a cgroup v2 Linux host, moving a PID into a cgroup that has controllers listed in its cgroup.subtree_control fails with EBUSY. What rule is the kernel enforcing, and how does a supervisor like systemd shape its tree around it?

level: seniorimportance: should knowfreq 30%

answer

  1. either processes or children, not both
  2. EBUSY from both directions
  3. the root is the exception
  4. controllers flow strictly downward
  5. policy on branches, processes on leaves

basics

~20 s

cgroup v2 forbids a non-root cgroup from both holding processes and distributing resources to its children — the no-internal-processes rule. Processes belong in leaves; inner cgroups only enable controllers via cgroup.subtree_control. That is why systemd puts services in leaf units under slices.

solid answer

~60 s

The kernel is enforcing the "no internal processes" constraint of the unified hierarchy: a non-root cgroup may either contain member processes or enable controllers for its children, never both. It bites in both directions — writing a PID into `cgroup.procs` of a cgroup whose `cgroup.subtree_control` is non-empty fails, and enabling a controller in `cgroup.subtree_control` of a cgroup that already has member processes fails too, both with `EBUSY`. The reason is that resource distribution would otherwise be ambiguous: a controller would have to arbitrate between child cgroups and loose processes sitting at the same level, which cgroup v1 did with confusing results. Only the root cgroup is exempt. Controllers also flow strictly downward: a cgroup's `cgroup.controllers` lists what its parent enabled for it, and it can only pass on a subset. Supervisors therefore build inner nodes that carry limits and leaves that carry processes — in systemd's model, slices are the inner nodes and services and scopes are the leaves, with `Delegate=` handing a subtree to a manager that wants to build its own tree beneath it.

code

bash · 10 lines
bash
# controllers must be enabled downward, one level at a time
cat /sys/fs/cgroup/cgroup.controllers
echo "+cpu +memory" > /sys/fs/cgroup/cgroup.subtree_control

mkdir -p /sys/fs/cgroup/app/worker
echo "+cpu +memory" > /sys/fs/cgroup/app/cgroup.subtree_control

# leaf accepts the process; the inner node returns EBUSY
echo $$ > /sys/fs/cgroup/app/worker/cgroup.procs
echo $$ > /sys/fs/cgroup/app/cgroup.procs

go deeper

for a junior

Know that a cgroup v2 process is placed by writing its PID into a cgroup.procs file, and that processes normally sit in the leaves of the tree rather than anywhere you like.

for a middle

State the no-internal-processes rule and both EBUSY directions, and explain the roles of cgroup.controllers versus cgroup.subtree_control in enabling controllers top-down.

for a senior

Explain why the constraint exists — unambiguous resource distribution — name the root exemption, and describe the inner-node/leaf shape it forces on any supervisor managing many workloads.

for a principal

Own the ownership question: who writes the cgroup tree on your hosts, which subtrees are delegated to other managers, and where resource policy is expressed so two managers never fight over the same nodes.

## What the rule says In the cgroup v2 unified hierarchy, mounted at `/sys/fs/cgroup`, a non-root cgroup may **either** contain processes **or** distribute resources to its children — not both. The kernel documentation calls this the "no internal process" constraint. It is enforced at both write points: - Writing a PID into `cgroup.procs` of a cgroup whose `cgroup.subtree_control` is non-empty → `EBUSY`. - Writing a controller into `cgroup.subtree_control` of a cgroup that already has member processes → `EBUSY`. The root cgroup is exempt, because there has to be somewhere for the processes that nobody has classified — kernel threads and early boot processes — to live. ## Why the kernel insists on it A controller distributes a resource among the entities immediately below a cgroup. If a cgroup could have both children *and* its own loose processes, the controller would have to decide how a process competes against a whole subtree. cgroup v1 allowed this and the semantics were both inconsistent between controllers and hard to reason about, so v2 removed the ambiguity by construction: at every level, a controller sees only cgroups. The practical effect is a shape rule. Your tree becomes inner nodes that carry policy and leaves that carry processes. ## The three files that matter - `cgroup.controllers` — read-only; lists the controllers this cgroup **may** use, because its parent enabled them. - `cgroup.subtree_control` — read-write; lists the controllers this cgroup passes **down to its children**. You may only enable something that appears in `cgroup.controllers`. - `cgroup.procs` — the member processes. Writing a PID moves that process here. Controllers are therefore enabled top-down, one level at a time: ```bash cat /sys/fs/cgroup/cgroup.controllers # cpuset cpu io memory hugetlb pids rdma misc echo "+cpu +memory" > /sys/fs/cgroup/cgroup.subtree_control mkdir -p /sys/fs/cgroup/app/worker echo "+cpu +memory" > /sys/fs/cgroup/app/cgroup.subtree_control echo $$ > /sys/fs/cgroup/app/worker/cgroup.procs # leaf: OK echo $$ > /sys/fs/cgroup/app/cgroup.procs # inner node: EBUSY ``` Disabling a controller in a parent's `subtree_control` also removes its interface files from the children, so limits set there are lost — enabling and disabling controllers is not a free operation on a live tree. ## The threaded exception There is one nuance worth knowing so you do not overstate the rule. A cgroup can be switched to *threaded* mode by writing `threaded` to its `cgroup.type`, which creates a threaded subtree where individual threads rather than whole processes are the unit of membership and `cgroup.threads` becomes the file you write. Inside such a subtree the constraint applies differently, because thread-granular controllers (the CPU controller in particular) must be able to arbitrate between threads of one process. Mentioning that you know the exception exists is usually enough; the domain-cgroup rule is what the question is really about. ## How supervisors build around it Because processes must live in leaves, any manager of many workloads ends up with a two-layer shape: grouping nodes that carry limits, and per-workload leaves that carry PIDs. systemd's cgroup layout is exactly this. Slices are the inner nodes — the root slice, and below it the slices for system services, for user sessions, and for machines — and services and scopes are the leaves where processes actually sit. Setting a limit on a slice constrains everything beneath it; setting one on a service constrains that unit. The second half of the picture is **delegation**. On a systemd host, systemd is the single writer of the cgroup tree, and two agents writing the same tree fight — one creates a cgroup, the other reaps it as unmanaged. When another manager legitimately needs to build its own subtree (a session manager, a container manager), it is *delegated* one: the unit declares `Delegate=`, and systemd hands over ownership of that subtree, changing the ownership of the directory, `cgroup.procs`, `cgroup.threads` and `cgroup.subtree_control` so an unprivileged manager can write them. The kernel's `nsdelegate` mount option reinforces that boundary by treating a cgroup namespace root as a delegation boundary that migrations may not cross. The operational lesson: on a systemd host, do not `mkdir` cgroups directly under the root and expect them to survive. Ask for a delegated subtree, or express the limit through the supervisor. ## What a good answer sounds like Name the constraint, show that you know it fires in both directions with `EBUSY`, explain *why* — resource distribution has to be unambiguous — note the root exemption, and land on the shape it forces: policy on inner nodes, processes in leaves, with delegation as the way two managers coexist.

  • Why is the root cgroup exempt from the no-internal-processes rule?
    Because something has to hold the processes nobody classified — kernel threads and everything running before any manager built a tree. Exempting the root keeps the system bootable, and it costs nothing semantically because the root is not competing with a sibling for a parent's budget.
  • What does reading a cgroup's cgroup.controllers file tell you, as opposed to cgroup.subtree_control?
    `cgroup.controllers` is read-only and lists what this cgroup *may* use, which is exactly what its parent enabled for it. `cgroup.subtree_control` is what this cgroup passes *down* to its own children, and you can only enable a subset of what `cgroup.controllers` shows. That is what makes controllers flow strictly top-down.
  • Why is creating cgroups by hand under /sys/fs/cgroup a bad idea on a systemd host?
    Because systemd treats itself as the single owner of that tree and will not preserve cgroups it did not create — yours can be removed underneath you, and limits are lost. The supported route is a delegated subtree, requested with `Delegate=` on a unit, after which the manager owning that subtree may build freely inside it.

saying these in an interview costs you the question

  • Thinks processes may live in any cgroup regardless of controllers
  • Says a child can enable a controller its parent did not
  • Believes EBUSY here means the cgroup is in use by another tool
  • Claims cgroup v1 enforced the same constraint
  • Creates cgroups directly under the root on a systemd host

context