skip to content

A batch job declares 500 required completions with 20 attempts running at once — what does the platform count, and when does it stop?

level: middleimportance: nice to knowfreq 33%

answer

  1. one number ends it, one widens it
  2. only clean finishes count
  3. top up to the smaller of the two
  4. the last wave is narrower
  5. a retry is budgeted, a replacement is not

basics

~20 s

It counts clean finishes, not started attempts. The at-once number is only a width limit: the platform keeps roughly that many attempts in flight, tops them up as each one ends, and marks the job finished when 500 have succeeded.

solid answer

~50 s

Two numbers do two different jobs. The **required completion count** is the end condition — the platform is finished when that many attempts have ended cleanly. The **at-once number** is a throttle on width, bounding how many attempts run in parallel so the batch does not take the whole cluster. The loop is: keep starting attempts until either the at-once bound or the outstanding work runs out, and on every clean finish increment the count and top up. Near the end fewer than 20 run, because there is less than that left outstanding. A failed attempt is not a completion — it spends a retry budget and, if the budget survives, another attempt for the same work is started. That retry is a different thing from a service's replacement: the job is accumulating successes toward a terminal state, while a service is holding presence and never accumulates anything.

code

pseudocode · 16 lines
pseudocode
on attemptEnded(job, endedCleanly):
    if endedCleanly:
        job.successes = job.successes + 1
        if job.successes >= job.requiredCompletions:
            markFinished(job)            # nothing more is started
            return
    else:
        job.attemptsSpent = job.attemptsSpent + 1
        if job.attemptsSpent >= job.retryBudget:
            markFailed(job)              # the platform gives up
            return

    outstanding = job.requiredCompletions - job.successes
    width       = min(job.runAtOnce, outstanding)
    while countRunning(job) < width:
        startAttempt(job)

go deeper

for a junior

Keep the two numbers apart: one says how much work must succeed, the other says how many attempts may run at the same time. Only clean finishes move the job toward being done.

for a middle

Walk the loop: on each ending attempt, count it or charge it to the retry budget, then top the running set back up to the smaller of the width bound and the work outstanding. Explain why the final wave is narrower.

for a senior

Bring in the properties the work must have — partitionable, idempotent, independent — and say what goes wrong without each. Then discuss choosing the width against the always-on workloads sharing the same hosts.

for a principal

Own the question of when a batch should be allowed to give up. A finite retry budget with a terminal failed state is a deliberate policy about how much of the night a broken job may consume before a human is told.

## Two numbers, two different jobs A run-to-completion workload that processes many items is declared with two independent numbers, and confusing them is the usual error. | Number | What it means | What happens if you change it | |---|---|---| | Required completions | **how much work must succeed** before the workload is done | the amount of work changes | | Attempts at once | **how wide** the work may run in parallel | the batch finishes sooner or later; the work is the same | The first is an end condition. The second is a throttle. A job with 500 required completions does the same amount of work whether 1 or 20 attempts run at a time — it simply takes twenty times longer at width 1. ## What the platform actually counts It counts **clean finishes**, and nothing else advances it: - an attempt that ends cleanly increments the completion count; - an attempt that ends badly does **not** — it is a failed attempt, charged against a retry budget; - an attempt that is still running counts toward the width bound but not toward completion. The loop the platform runs is small. On each attempt ending, update the counters, check the terminal conditions, then top the running set back up to the smaller of the at-once bound and the work still outstanding. ## Working the stated example Start: 500 required, 20 at once, 0 succeeded. The platform starts 20 attempts. As each one finishes cleanly the count rises and a replacement attempt starts, so the width stays at 20 while at least 20 units of work remain outstanding. Once only 12 remain outstanding, at most 12 attempts run — the width bound stops being the constraint and the remaining work becomes the constraint. When the 500th clean finish lands, the workload becomes **finished** and nothing further is started, even though moments earlier there were attempts in flight. If 5 attempts fail along the way, those 5 do not count. The platform starts fresh attempts for that work while the retry budget allows; if the budget runs out the workload becomes **failed**, a terminal state that stops new attempts even though fewer than 500 completions were reached. ## A retry is not a restart Both words describe "it runs again", which is why this distinction is worth stating precisely. | | Retry, under the job contract | Replacement, under the service contract | |---|---|---| | Why it happens | an attempt ended badly and work is still outstanding | the number of live copies is short | | What it is counted against | a finite retry budget | nothing — it is unbounded | | Terminal state | yes: finished or failed | none; the workload runs until you change it | | What it means for the work | the same unit of work is attempted again | no unit of work is implied at all | The practical consequence is that a job can **give up**, and a service cannot. That is a feature: a batch whose work is genuinely broken should stop and say so before morning, rather than burning the night on it. ## What the work items have to be Parallel attempts and retries only make sense if the work is shaped for them: 1. **Partitionable** — 500 completions implies 500 separable units, and something must assign an attempt to a unit deterministically or hand units out from a shared queue. 2. **Idempotent** — a retried attempt may be re-doing work a previous attempt partly finished, so applying it twice must land in the same state as applying it once. 3. **Independent** — attempts running at once must not depend on each other's order, or the width number silently becomes a correctness setting instead of a throughput setting. When those three hold, the at-once number is a pure speed-against-load dial: raise it to finish the nightly batch inside its window, lower it when it starves the always-on services sharing the same hosts. When they do not hold, raising it corrupts data, and no platform setting will tell you so.

  • Why does the last wave of a wide batch run narrower than the declared at-once number?
    Because the width is the smaller of the at-once bound and the work still outstanding. With 500 required and 20 at once, the bound is what limits you until fewer than 20 units remain; after that the remaining work does. Starting more attempts than there is outstanding work would either duplicate a unit or run an attempt with nothing to do.
  • What breaks if the units of work are not idempotent?
    Retries stop being safe. A failed attempt may have applied part of its effect before ending, so the replacement attempt re-applies it — double charges, double rows, double sends. Nothing in the contract detects this: the platform only knows that an attempt ended badly and another is allowed. Idempotence is the property that makes the retry budget usable at all.
  • Is raising the at-once number ever unsafe rather than merely faster?
    Yes, in two ways. If the units are not independent, width turns into a correctness setting and concurrent attempts collide. And even when they are independent, a wide batch competes for host capacity with the always-on services beside it, so the batch finishing sooner can mean the services serving worse. Width is a dial with two costs, not just a speed control.

saying these in an interview costs you the question

  • Says the at-once number decides how much work the job does
  • Counts started attempts rather than clean finishes toward completion
  • Thinks a failed attempt still moves the job closer to finished
  • Claims a job keeps retrying forever, like a service being replaced
  • Assumes the declared width runs right up to the final attempt
  • Treats retries as safe without the work being idempotent