A search indexer that needs four minutes to load its index is killed at the same point on every attempt — why?
answer
- same elapsed time on every attempt
- counting starts before serving can
- start-up needs its own allowance
- hold the restarting verdict off first
- allowance sized above the worst cold start
basics
~20 sThe liveness check starts counting before the index is loaded, so the same failure count is reached at the same elapsed time on every attempt and the platform restarts it — forever. The fix is a start-up allowance, not a looser restart threshold.
solid answer
~50 sThe restarting verdict has no concept of *has not served yet*. Its check begins as soon as the container is running; the indexer cannot answer it while it is loading, so consecutive failures accumulate, the failure threshold is reached, and the platform ends the process — always at the same elapsed time, because nothing about the timing varies between attempts. The next attempt starts the same clock and dies at the same point, so the indexer never reaches the four minutes it needs. The correct fix is a **startup check**: a bounded allowance that holds the liveness and readiness verdicts off until start-up reports itself finished, or until the allowance runs out. Raising the liveness failure threshold to four minutes also stops the kills, but it dulls detection of a genuine hang for the entire life of every instance.
code
yaml · 22 lines# one workload's three declared checks; all times in seconds
checks:
startUp: # runs first; holds the other two off
endpoint: /started
interval: 10
timeout: 2
maxFailures: 40 # allows up to 400s of start-up, then gives up
liveness: # begins only after startUp first succeeds
endpoint: /alive
interval: 10
timeout: 2
maxFailures: 3 # verdict after ~30s of continuous failure
onVerdict: restart
readiness: # decides routing only; never ends the process
endpoint: /canServe
interval: 5
timeout: 2
maxFailures: 2
onVerdict: leaveRoutingSetgo deeper
Recall that a restarting check begins immediately unless something holds it off, and that a slow start looks exactly like a hang to it. Name the start-up allowance as the fix.
Explain the determinism: the failure threshold times the interval gives the same verdict at the same offset on every attempt, which is why the instance never reaches the four minutes it needs.
Show that you would size the allowance against the worst observed cold start rather than the median, and argue why a blunt threshold increase is a regression disguised as a fix.
Treat start-up cost as a fleet property: the slower a workload is to become useful, the more expensive every restart verdict is, and the more the platform's defaults need to be set deliberately rather than inherited.
## The symptom is most of the diagnosis The distinguishing feature here is not that the instance dies — it is that it dies at *the same point every time*. Deterministic timing rules out most causes. A process failing on its own input, on a corrupt file or on a missing dependency fails at a point determined by its work, which varies with data and node speed. A death at a fixed wall-clock offset, repeated identically, points at something outside the process counting seconds: a health verdict. Two more observations confirm it: - the **previous attempt's output** stops mid-load with no error and no shutdown line — the process was ended, it did not decide to exit; - the elapsed time before the kill matches `interval x failureThresholdCount`, not anything in the start-up work. ## Why the restarting check cannot tell "starting" from "stuck" A liveness check only ever sees success or failure. An instance four minutes into loading an index and an instance permanently deadlocked produce the identical answer: no success. The check has no way to distinguish them, and it was never given one — so if it is allowed to run during start-up with ordinary settings, a slow start is indistinguishable from a hang and is punished as one. The loop is closed and self-sustaining: kill, restart, same clock, same verdict, kill. The instance never survives long enough to serve a single query, and the repeated restarting that follows is only the visible symptom of a verdict that was wrong the first time. ## Three jobs, and the one that is missing | check | what its failure does | what it is for | |---|---|---| | startup | nothing to routing, nothing to the process while it runs; on exhaustion the instance is given up on | buying a bounded start-up window | | liveness | ends the process and starts a fresh one | a process that is stuck and will not recover | | readiness | removes the instance from the routing set, leaves it running | an instance that is alive but cannot serve right now | The indexer needs all three, and the missing one is the first. A **startup check** runs while the instance is starting and holds the other two verdicts off. Its threshold is deliberately generous, because its only job is to answer *is start-up still making progress, and has it taken longer than we are prepared to wait?* When it first succeeds, the allowance ends and the normal liveness and readiness settings take over for the rest of the instance's life. ## Sizing the allowance 1. **Measure the worst honest cold start**, not the median: the largest index, the coldest page cache, the slowest node, the busiest moment for whatever the instance reads from. 2. **Add real margin** — the allowance is a ceiling, not a target. An instance that finishes in ninety seconds moves on at ninety seconds; nothing is wasted by a generous cap. 3. **Set the cap where you would genuinely give up.** Past the allowance, the platform stops waiting, and that is the correct behaviour for a start-up that is truly wedged. 4. **Re-check it when the data grows.** An allowance sized against last year's index is the same defect waiting to return. ## Why a raised failure threshold is the worse fix Raising the liveness threshold until it exceeds four minutes does stop the kills. It also applies for the entire life of every instance: a genuinely hung instance now keeps running, and keeps its place in the fleet's capacity, for minutes before anything acts. The whole point of a start-up allowance is that it is **spent once**, so patience during start-up costs nothing in steady state. A per-attempt **timeout** is not a fix either, and reaching for it is a common wrong turn: a timeout bounds how long *one* check attempt may take, not how many attempts may fail. Making it ten times larger changes nothing about a start-up that cannot answer at all. ## Readiness during start-up While the index is loading, the instance should also be *unready*: it is running, but it cannot serve. Leaving it in the routing set means traffic reaches an instance that will fail or stall every request. Readiness is the cheap, reversible way to say so, and unlike the restarting verdict it costs nothing to be wrong about for a few seconds.
- Why is raising the liveness failure threshold a worse fix than adding a start-up allowance?Because the raised threshold applies forever. After start-up, a genuinely hung instance is tolerated for as long as the threshold now permits, so you have traded a start-up bug for permanently blunt detection. An allowance is spent once and the strict settings resume.
- What does the platform do if the start-up allowance runs out before start-up finishes?It stops waiting and treats the instance as failed, which is the intended behaviour — the allowance is a cap on patience, not an unlimited exemption. That is why the cap should sit where you would genuinely give up rather than at an arbitrary round number.
- From outside, how do you tell a missing start-up allowance from a process that is crashing on its own?By the determinism and by who ended it. A crash lands at a point set by the work and usually leaves an error and a non-zero exit of its own. A missing allowance ends the instance at a fixed elapsed time matching the check's interval times its threshold, with the previous attempt's output stopping mid-load and saying nothing.
saying these in an interview costs you the question
- Blames the image or the index data rather than the check
- Assumes the platform waits for start-up before checking anything
- Claims a longer per-attempt timeout would let start-up finish
- Raises the restart threshold so high a real hang is never caught
- Says the restarts happen at random points rather than a fixed one
- Thinks a failing readiness check is what ended the process