skip to content

A long-running queue consumer was declared under the run-to-completion contract and exits cleanly when the queue empties — what does the platform do?

level: middleimportance: must knowfreq 58%

answer

  1. nothing was violated
  2. a count of runs, not of copies
  3. a clean exit is the declared goal
  4. the platform is green, the backlog is not
  5. presence invariant is what you actually wanted

basics

~20 s

Nothing. A clean exit is what the run-to-completion contract was told to expect, so the platform records a successful run and stops. The consumer is gone, no copy is missing, nothing failed, and no alert fires.

solid answer

~40 s

The run-to-completion contract holds a count of successful runs, not a count of live copies — so there is no running-copy invariant to violate when the process leaves. The platform increments completions, sees the target met, marks the workload finished and keeps that record. From the platform's side everything is green: no restart, no failure, no missing copy. What actually happens is that message backlog grows and downstream work stops, and you find it through the queue's own metrics rather than through the platform. The fix is to declare it as a replicated long-running service so the invariant becomes `this many copies exist`, at which point any exit — clean or not — is a shortfall the platform fills and a repeatedly exiting consumer becomes loudly visible instead of silently absent.

go deeper

for a junior

Remember that a run-to-completion workload is allowed to end. If a process that should have kept running exits cleanly under that contract, the platform treats it as work well done and nothing replaces it.

for a middle

Explain it from the invariant: a count of successful runs was satisfied, so no shortfall exists to reconcile. Then contrast it with the presence invariant a replicated service holds, and say which exits each one reacts to.

for a senior

Show how you would have detected it. Argue for alerting on the work — backlog age, output freshness — because a platform whose contract was satisfied will report nothing, and describe how you would re-declare the workload without touching the program.

for a principal

Treat it as a defaults question. Decide which contract a new workload gets by default, what review catches a continuous consumer declared to finish, and how much of your alerting is allowed to depend on the platform's own view of health.

## What the contract actually promised A run-to-completion workload is a promise about **outcomes**, not about **presence**. The declaration says: run this work until a declared number of runs have finished successfully, then stop. The platform's loop for it therefore asks "have enough runs succeeded yet?" — and when a process ends cleanly, the honest answer becomes yes. A consumer that drains a queue and returns when the queue is empty satisfies that literally. It ran. It ended without error. The completion count is met. The platform writes `finished: succeeded`, keeps the record, and has nothing left to do. There is no shortfall, because at no point did anyone declare that a copy should exist. ## Why this failure is quiet, and the service contract's is loud | | Declared as a job (what happened) | Declared as a service (what you wanted) | |---|---|---| | Invariant held | a count of successful runs | a count of live copies | | Clean exit | counts as success; workload finishes | a shortfall; another copy is created | | Unclean exit | a failed attempt against a retry budget | also a shortfall; another copy is created | | After the last exit | no copy exists, nothing is wrong | the platform keeps re-creating copies | | What you see | a green record and silence | repeated re-creation — obvious in the workload's own state | The asymmetry is the whole lesson. The wrong contract in the *other* direction — a batch declared as an always-on service — is noisy and self-announcing: the work runs again every time it finishes, and somebody notices by morning. The wrong contract in *this* direction is the dangerous one precisely because the platform has nothing to complain about. ## How you actually find it The platform is not the detector here; the work is. - **Backlog depth** on whatever the consumer reads climbs and does not come down. - **Downstream freshness** ages: records that should be minutes old are hours old. - **The workload's own record** shows a finished run with a suspiciously short duration and a completion time that matches when the trouble started. If your alerting is built only on "is the platform unhappy", this class of failure is invisible to it by construction. This is one of the clearest arguments for alerting on the *work* — queue age, output freshness, records processed per minute — rather than on the shape of what is running. ## The fix, and the case where the job contract was right If the consumer is meant to live forever, declare it as a replicated long-running service. Now the invariant is presence: any exit, clean or not, leaves the running count short and the platform creates another copy. A consumer that returns the instant its queue is empty will then be re-created continuously, which is not elegant — but it is *visible*, and you fix the process so its main loop waits for work instead of returning. There is a legitimate design where the job contract is the right one: a drain-this-batch-and-stop worker, started by a recurring schedule, deliberately processing whatever accumulated since the last run. The difference is not the code, it is the intent — you accept that nothing is consuming between runs, and you size the schedule against how stale the work is allowed to get. What makes the failure a failure is declaring the job contract for a workload whose requirement is continuous consumption. ## The direction that trips candidates Be explicit about which way the mistake runs, because both sentences are grammatical: 1. **Service declared as a job** — clean exit reads as success, nothing replaces it, the workload silently disappears. 2. **Job declared as a service** — clean exit reads as a missing copy, so the finished work is started again, and again. An interviewer asking this question is usually checking whether you can state the invariant each contract holds and then derive the symptom, rather than recalling an anecdote. The derivation is short: name the invariant, ask what a clean exit does to it, and the behaviour follows.

  • What signal would have caught this within minutes?
    A signal taken from the work rather than from the platform: the age of the oldest unconsumed message, or the freshness of what the consumer produces. Both go bad immediately when consumption stops, and neither depends on the platform considering anything wrong. Platform-side signals are useless here because, under the run-to-completion contract, nothing is wrong.
  • What is the symptom if you make the mistake the other way round — a finite batch declared as an always-on service?
    The opposite and much louder one. The batch finishes, the platform sees the declared copy count unmet and creates a copy, which does the same work again. You get the job re-run continuously instead of once, which shows up as duplicated output, wasted capacity and an obviously churning workload — a failure that announces itself by morning.
  • Is a consumer that exits when its queue is empty ever the right design?
    Yes, when the requirement is periodic drainage rather than continuous consumption: a recurring schedule starts a run, the run processes what accumulated and stops. You accept a gap between runs and pick the interval from how stale the work may get. It stops being right the moment someone expects messages to be handled as they arrive.

saying these in an interview costs you the question

  • Says the platform will restart it because every workload gets restarted
  • Thinks a clean exit always means the workload stops and an error always means it restarts
  • Claims the missing consumer will show up in platform health checks
  • Confuses this with the noisy case where a finished batch is started again forever
  • Assumes changing the contract requires changing the program's code
  • Believes a finished record with a short duration proves the run did its work