skip to content

A batch group's work container exits successfully but its metrics helper keeps running — why does the group never finish?

level: seniorimportance: nice to knowfreq 28%

answer

  1. the group finishes, not the member
  2. a helper with no exit condition
  3. one running member keeps the unit alive
  4. tie the helper's lifetime to the work
  5. or keep it out of the group entirely

basics

~20 s

Completion is a property of the group, not of one member. A run-to-completion group is done only once its members have stopped, so a helper with no exit condition keeps the group alive and the run never records success.

solid answer

~50 s

The work member finishing is not the group finishing. A group meant to run to completion is judged on all of its members, so a helper written to run forever — a collector, a refresher, a tailer — never lets the group reach a terminal state. It goes on holding its place on the host, whatever was waiting for the run keeps waiting, and the attempt is eventually cut off by a deadline instead of being recorded as a success. There are three honest fixes: declare the helper as a member whose lifetime is tied to the work member, so the platform stops it when that member exits (platforms differ in whether they offer this); have the work member signal the helper before exiting, through the shared scratch area or a call on the group's own address; or keep a never-exiting helper out of a run-to-completion group entirely.

code

pseudocode · 11 lines
pseudocode
groupFinished = true
for each member in group:
    if member.stillRunning:
        groupFinished = false

if groupFinished and every member exitedWith(0):
    record run as succeeded

# the metrics helper has no exit condition, so its stillRunning stays
# true forever, groupFinished never becomes true, and the run is never
# recorded as succeeded

go deeper

for a junior

Know that a group meant to finish is finished only when all of its members have stopped, not when the main one has.

for a middle

Explain the mechanism: completion is evaluated over the whole group, so a single member with no exit condition is enough to keep it running forever.

for a senior

Diagnose it from the symptoms — a work member with a clean exit, a group still holding host capacity, and a run that ends on a deadline instead of a success.

for a principal

Make it a rule for shared components: anything admitted into a group that has to finish must declare how it ends, or the platform must be able to end it for you.

## Completion is group-granular Everything else about a co-located group is decided for the group as a whole — where it is placed, when it is replaced, when it is deleted — and completion is no exception. A group whose purpose is to finish is finished when its members have stopped, not when the member you happen to care about has stopped. That single sentence is the whole answer, and it is counter-intuitive precisely because the work member is the one you were watching. So a helper that was written the way helpers are normally written — start, loop forever, be stopped by something else — holds the group open indefinitely once you place it alongside finite work. Nothing is broken; the helper is doing exactly what it was built to do. The mismatch is in the shape of the group. ## Why the helper does not stop on its own A helper container has no notion that its neighbour has finished. It has no signal from the platform saying 'the other member exited', and in most designs it is not watching the group's state at all — it is scraping, refreshing or tailing on a timer. Left alone it runs until something stops it, and in a long-running group that something is the group's own replacement or deletion. In a finite group, that something never arrives, because the group is waiting for the helper. ## What the stuck group costs - **Host capacity.** The group's reservation is still subtracted from the host it sits on, for as long as it sits there, even though no useful work is happening. - **Whatever waits on the run.** Anything gated on the run finishing is gated on a group that will not finish. - **The success record.** When a deadline eventually fires, the run ends as cut off rather than successful, so the outcome recorded is wrong as well as late. - **Attempt accounting.** Because the attempt never terminates cleanly, the bookkeeping the platform keeps about attempts does not advance either — the run is neither done nor retried. ## Diagnosing it 1. Confirm the work member actually exited, and with a success code. If it did, the workload is not the problem. 2. List the group's members and their states. One still running is the answer, and it will normally be the helper. 3. Check the helper's design for an exit condition. If there is none, this is a shape problem, not a bug in either member. The tell is a group that holds capacity with a work member that has been finished for hours. ## The three fixes, and what each costs | Fix | What it costs | |---|---| | Tie the helper's lifetime to the work member, so the platform stops it on exit | Depends on the platform offering it; designs genuinely differ here | | Have the work member tell the helper to exit before it does | A small coordination step you have to write and get right | | Keep the never-exiting helper out of the group | You give up the shared address and the shared scratch area | The second fix uses nothing but what the group already gives you: write a sentinel file into the shared scratch area that the helper polls, or call a shutdown endpoint the helper exposes on the group's shared address. Both are local, both are ordinary, and neither needs a platform feature. A **deadline** on the run is worth having regardless — it turns an unbounded hang into a bounded failure, which is much easier to alert on. But it is a guard, not the fix. It ends the run as cut off rather than successful, and it burns the whole deadline every time before doing so. ## The shape rule When a group has to finish, every member of it has to be able to finish. That is the rule to apply at the point of adding a member: a component with no exit condition belongs in a group that is not expected to end, or has to be given one — either by the platform stopping it, or by your own code telling it to go. Co-locating an endless process with finite work and expecting the platform to sort it out is how you get a run that is neither successful nor failed.

  • Why not simply put a deadline on the group?
    A deadline turns an unbounded hang into a bounded failure, which is worth having as a guard. But it records the run as cut off rather than successful, and it burns the full deadline on every run before doing so. It does not fix a member that was never given an exit condition.
  • How can the work member tell the helper to stop?
    Through what the group already shares: write a sentinel file into the shared scratch area that the helper polls, or call a shutdown endpoint the helper exposes on the group's address. Both keep the coordination inside the group and need no platform feature at all.
  • Does this ever happen in a long-running group?
    No, and that is why it catches people out. In a group that is never meant to end, a helper that never exits is exactly right, so the same helper image is reused in a finite group without anyone thinking about it. The defect only appears when completion starts to matter.

saying these in an interview costs you the question

  • Assuming the group ends when the main member exits
  • Expecting the platform to stop the other members automatically
  • Treating a stuck run as a slow job
  • Adding a deadline and calling the cause fixed
  • Co-locating a never-exiting helper with finite work