skip to content

A workload instance has been alive for ten minutes, was never replaced, and still receives no traffic — which startup failure is it?

level: middleimportance: should knowfreq 58%

answer

  1. alive, unreplaced, and unused
  2. uptime climbing, not resetting
  3. no termination record at all
  4. routing is opt-in on a ready report
  5. still working, blocked, or mismatched

basics

~20 s

The fourth one: the process runs but has never reported itself ready, so the platform keeps it out of routing. A climbing uptime with no replacement and no termination record rules out a process that started and died.

solid answer

~50 s

This is the only one of the four startup failures where the instance itself is fine. It was fetched, it was started, the process is alive and its output is flowing — and the platform has no positive ready report for it, so it sends nothing its way. That is the designed behaviour, not a routing bug. The evidence that separates it from a process that starts and dies is all in one direction: uptime climbing rather than resetting, no termination record, no replacement, and your own output in the stream. Having named the stage, the next question is whether the process is still finishing start-up work, whether it is blocked on something it cannot reach, or whether the readiness check is examining something the process never satisfies — a separate subject with its own mechanics.

go deeper

for a junior

Remember that being alive and being ready are different things. Traffic arrives only after the instance reports it can serve, so an instance that is up and idle has not made that report.

for a middle

Explain each discriminator against a process that starts and dies — uptime climbing, no termination record, no replacement — and say what the platform does instead of tearing the instance down.

for a senior

Show the triage split between still working, blocked and a reporting mismatch, and explain why replacing the instance destroys the only live evidence while fixing nothing.

for a principal

Argue for making this failure loud: it is the quiet one, it holds a place indefinitely, and anything waiting on readiness stalls with no error for someone to be paged about.

## Why this one looks like success Three of the four startup failures announce themselves. Something is missing, something is being retried, something died. This one presents as a live process with a growing uptime, ordinary start-up lines in its output and nothing being replaced. The only symptom is negative: no requests arrive. That shape is exactly what the platform intends. Routing to an instance is opt-in: the platform adds an instance to the set traffic is sent to only once it has a positive ready report for it. No report means no traffic, indefinitely and without an error. Nothing in the picture is broken in the sense the other three stages are broken. ## Separating it from a process that starts and dies The two are confused constantly, because both mean "the workload is not serving". Every discriminator points one way: | Observation | Started and died | Runs and never reports ready | |---|---|---| | Uptime of the current instance | resets repeatedly | climbs continuously | | Termination record | present | absent | | Replacement of the instance | happening | none | | Output of the process | often stops mid-start-up | continues normally | | In routing | no | no | Only the last row is shared, and it is the one that produced the page. Reading any of the other four settles it in seconds. ## What the platform is and is not doing - It **leaves the process running**. A missing ready report withholds traffic; it does not tear the instance down. - It **keeps the instance's place**. The workload's slot is occupied by something that will never serve, which is why this failure can be expensive despite looking calm. - It **reports nothing as an error**, because from its point of view nothing has gone wrong yet — it is still waiting. - It **stalls whatever is waiting on readiness**. A replacement that waits for the new instance to be ready before continuing will sit there, and the absence of an error is exactly why nobody notices for an hour. ## Where the cause usually lives, at triage altitude Once the stage is named, there are three families of cause worth separating, and you can usually tell them apart from the process's own output: 1. **Still working.** The process is genuinely not ready — loading a large dataset, warming a cache, waiting for a schema step to finish. The output is progressing. The question becomes whether the platform was ever told to expect a long start. 2. **Blocked.** The process cannot complete start-up because something it needs is unreachable, and it is looping or waiting quietly. The output usually shows a repeating attempt, or stops at the same line every time. 3. **Reporting mismatch.** The process is serving perfectly well and whatever asks it for a ready answer never gets a positive one. The output shows a normal, finished start-up. The mechanics of that check — what it asks, how often, how many failures count — are a separate subject; triage's job is to have ruled out the first two. A fourth possibility is worth naming only to dismiss it: an instance that *was* ready and then dropped out of routing is not this failure at all. That one has a period of served traffic behind it, and it belongs to a running workload's behaviour rather than to startup triage. ## The trap The reflex, when a live instance will not serve, is to replace it. That destroys the only live example of the failure you have — a process sitting in exactly the state you needed to inspect — and, because the instance was never the problem, the replacement arrives in the same state a minute later. Two things are worth doing before replacing anything: read the current instance's output to the end and decide which of the three families it is in, and record how long the instance has been alive, because the answer to "was it ever going to become ready, given more time?" is only available while it is still there. The opposite reflex is just as costly: declaring the instance healthy because it is up. Uptime is not readiness. An instance that is alive and out of routing is contributing nothing at all, and if the platform is waiting on it, it is also holding something else still.

  • How do you tell "never became ready" from "was ready and then dropped out of routing"?
    By whether the instance ever served. A never-ready instance has no traffic anywhere in its history and an uptime that covers the whole window. One that dropped out has a period of served requests followed by removal, with the process usually still alive. Only the first is a startup failure.
  • Why is this failure more expensive than a process that dies at once?
    Because it is quiet. A dying process produces terminations, replacements and error output that something will alert on. A never-ready instance produces none of those: it holds its place, serves nothing, and anything waiting for it to become ready waits without an error to report.
  • What do you lose by replacing the instance immediately?
    The only live copy of the failure. The process is sitting in the exact state you want to inspect, and the replacement will reach the same state, so nothing is gained. Read its output to the end and note how long it has been alive first.

A shop with the lights on, staff inside and the till open, but the sign on the door still reads closed. Nothing inside is broken — nobody flipped the sign, so the street walks past.

saying these in an interview costs you the question

  • Calls a running, unready instance a crash loop
  • Assumes no traffic means the process must be dead
  • Declares it healthy because the uptime is climbing
  • Expects routing as soon as the process is listening
  • Replaces the instance and loses the live evidence
  • Blames the routing layer before checking the ready report