A host's container management service was restarted, yet every workload kept running and kept its output stream — how?
answer
- the manager is not the parent
- something small sits in between
- who holds the pipes and the exit status
- the creating runtime already exited
- bundle and state on disk, re-attach later
basics
~20 sThe manager is not the workload's parent. The low-level runtime that created each container exited at start-up, and a small per-container supervising process sits between: it parents the first process, owns its output pipes and holds its exit status until the manager returns and re-attaches.
solid answer
~40 sNothing in the running container's lineage belongs to the manager. The low-level runtime built the boundary, executed the declared command and exited; what the manager left behind is a **supervising process**, one per container, running on the host outside the boundary. It is the parent of the container's first process, it holds the write ends of the standard output and error pipes and forwards them to wherever logs are collected, and it waits for the process so the exit status is captured whether or not the manager is up. Each container's bundle and state record live on disk, so a restarted manager re-discovers what is running and re-attaches to each supervising process. The consequence engineers actually use: the container stack can be upgraded under load without draining the host.
code
pseudocode · 15 lines# host process tree, minutes after start-up
container-manager # restartable; holds no pipes, no child workloads
supervising-process-A # child of the manager, OUTSIDE the boundary
ingester-process-A # first process INSIDE the boundary
supervising-process-B
batch-worker-B
# low-level-runtime: not present -- it exited right after it
# applied the fences and executed the command
on manager_restart:
for record in disk.state_records(): # bundle path + container id + channel
reattach(record.supervisor_channel) # a lookup, not a restart
# workloads never noticedgo deeper
The takeaway to recall is that the service you interact with is not the parent of your process. Restarting it does not stop workloads, because something smaller was left holding each container.
Explain the three host-side actors and their relationship: a restartable manager, a per-container supervisor that owns the output pipes and waits for the exit status, and the first process inside the boundary. Note that the creating runtime already exited.
Draw the operational conclusion and its limits: upgrading the container stack under load is safe and draining is unnecessary, but the supervisor is a single point of failure for one container's logs and exit status, and nothing here survives a kernel-level failure.
Frame it as a design rule for node-level agents generally: keep anything holding a file descriptor or a pending status small enough that it never needs upgrading, and keep rebuildable state on disk. That is what makes a fleet's agent rollout a non-event.
This is one of the better questions on this material because the naive model — a daemon that owns everything it started — predicts exactly the opposite of what the host does. Working through why the prediction fails is working through the layering. ## Who is whose parent After a container is running there are three distinct things on the host, and only one of them is inside the boundary: - **The container manager**, a restartable service that resolved the image, unpacked it and wrote the configuration document. It is the component being upgraded, and it is *not* in the workload's process lineage. - **A supervising process**, one per container, started by the manager and left behind deliberately. It runs on the host, outside every fence the container got. - **The workload's first process**, inside the boundary — the command the configuration document declared. The low-level runtime that created the boundary is absent from that list, because in the common design it exited the moment it executed the declared command. ## What the supervising process is holding It is deliberately tiny, and it exists to hold exactly the things that must not be lost when the large component restarts: 1. **The output streams.** It owns the read ends of the container's standard output and error and writes them onward — to a file the platform rotates, or to whatever collects logs on the host. Because the pipes live in a process nobody restarts, a manager upgrade does not truncate or lose the stream. 2. **The exit status.** It is the parent, so it is what the kernel notifies when the first process ends. It waits, records the status and holds it until something asks. A container that exits while the manager is down still reports its exit code correctly afterwards. 3. **The terminal, when one was allocated.** An interactive session survives the same way. ## Why the state survives a restart The other half of the answer is that the manager keeps almost nothing in memory that it cannot rebuild. Each container's bundle — the unpacked root filesystem and the configuration document — is on disk, and so is a small state record naming the container, its identifier and how to reach its supervising process. On start-up the manager reads those records, re-attaches over each container's own channel and resumes reporting on containers it did not start in this lifetime. ## What this buys, and what it does not | Event | What happens to running workloads | |---|---| | Container manager restarted or upgraded | Keep running; output and exit status preserved; the manager re-attaches | | Low-level runtime binary replaced on disk | Running containers unaffected — their runtime already exited | | A supervising process is killed | Workload keeps running, but its output collection and exit status are lost | | Host kernel panic or reboot | Everything on the host stops; this layering does not help | The limits deserve honesty. The supervising process is a single point of failure for one container's observability, not for its execution: kill it and the first process keeps serving, is re-parented on the host, and its ending is simply never reported. And designs genuinely differ — some stacks keep the low-level runtime resident as the container's parent instead of leaving a separate supervisor, which changes what a restart of the layer above costs. The property to check on any stack is the same one: *is the component I am about to restart in the workload's process lineage, and does it hold any pipe or status I cannot rebuild?* ## How to use it in an interview The strong version of this answer resists the temptation to say "containers are independent of the daemon" as a slogan and instead names the mechanism: the creating runtime already exited, a small supervisor holds the pipes and the exit status, the bundle and state are on disk, and re-attachment is a lookup rather than a restart. It also draws the correct operational conclusion — upgrading the container stack on a node is not a reason to drain it, while anything that touches the kernel is.
- What is actually lost if the supervising process itself is killed?The container's observability, not its execution. The first process keeps serving and is re-parented on the host, but the pipes it was writing into are gone, so log collection stops, and nothing is left waiting for it, so its eventual exit status is never recorded. The platform will typically report the container in an unknown or stopped state while the process is still doing work.
- Why not let the manager hold each container's pipes and exit status directly?Because then every restart of the biggest, most frequently upgraded component would close those pipes and discard the exit status of anything that ended while it was down. Splitting out a process small enough never to need upgrading is what makes the rest of the stack safely restartable under load.
A night porter holds your mail and your room key while the front desk changes shift. The desk can be replaced entirely; your mail is still there, and the new shift just asks the porter.
saying these in an interview costs you the question
- Assumes every container dies when its manager restarts
- Thinks the manager itself holds each workload's output pipes
- Believes the low-level runtime stays resident supervising the workload
- Says the exit status is always lost if the manager was down
- Treats the supervising process as the container's own first process