What problem does the Kubernetes startupProbe solve, and how does it interact with the liveness and readiness probes while it is still running?
answer
- slow boot vs fast hang detection - one knob, two needs
- while pending: liveness and readiness disabled
- budget = failureThreshold x periodSeconds
- succeeds once, then never runs again
- lets liveness be aggressive
basics
~20 sIt protects slow-starting containers: while a startupProbe is defined and has not yet succeeded, the kubelet suspends the liveness and readiness probes and the container is not ready. Once it succeeds it never runs again. Its budget is failureThreshold x periodSeconds.
solid answer
~50 sBefore startup probes you faced a bad trade-off: either give the liveness probe a large `initialDelaySeconds`/`failureThreshold` - which also blunts detection of real hangs forever after - or watch the container be killed mid-boot and crash-loop. `startupProbe` splits the two phases. While it is defined and not yet successful: - the **liveness** probe does not run, so the container cannot be restarted for being slow; - the **readiness** probe does not run either, and the container is treated as not ready, so no traffic arrives. Its total allowance is `failureThreshold x periodSeconds` (plus `initialDelaySeconds`); exceed it and the container is killed and restarted per `restartPolicy`. On first success it is **never executed again**, and the normal liveness/readiness cadence takes over - so those can now be tight and fast-reacting. Typical shape: a startup probe with a generous budget (30 x 10s = 5 minutes) for a JVM or a service replaying large state, plus a liveness probe with period 10 and failureThreshold 3.
code
yaml · 20 linescontainers:
- name: legacy-app
image: registry.example.com/legacy:9.1
startupProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 10
failureThreshold: 30
livenessProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: 8080
periodSeconds: 5go deeper
Say it gives slow-starting containers time to boot without being restarted, and that liveness takes over once it passes.
Add the exact interactions - liveness and readiness suppressed while pending, runs only until first success - and the failureThreshold x periodSeconds budget.
Explain why it lets you tighten liveness detection, how to size the budget from worst-case cold starts, and how to spot a too-small budget in restart events and logs.
Standardise startup budgets per workload class in platform templates, and reason about cold-start variance under node pressure, image pull time and dependency warm-up when setting fleet defaults.
## The problem Some containers take a long and *variable* time to become useful: a JVM warming JIT and connection pools, an app replaying a changelog or rebuilding an index, a legacy application server. A liveness probe that starts checking too early kills the container mid-boot; it restarts, boots slowly again, and you get a permanent CrashLoopBackOff that looks like an application bug but is really probe misconfiguration. Before Kubernetes 1.16 the only lever was to make the liveness probe itself tolerant - a huge `initialDelaySeconds` or `failureThreshold`. That works during boot and ruins steady state: afterwards, the same tolerance means a genuinely hung process takes minutes to be detected. One knob was being asked to serve two opposite requirements. ## What startupProbe does `startupProbe` is a third probe with the same handler options (`exec`, `httpGet`, `tcpSocket`, `grpc`) and the same timing fields. Its role is to answer one question once: *has this container finished starting?* Semantics while it is pending: - **Liveness is disabled.** The kubelet will not restart the container for liveness reasons, however long boot takes, as long as the startup probe has budget left. - **Readiness is disabled and the container is not ready**, so it stays out of Service endpoints. There is no window in which a half-started Pod receives traffic. - The startup probe itself retries every `periodSeconds` up to `failureThreshold` times. When it **succeeds**, it never runs again for that container instance, and liveness and readiness begin their normal cadence immediately. If it **exhausts** its budget, the kubelet kills the container and `restartPolicy` applies - usually another restart, with the budget starting over. ## Sizing it The maximum tolerated startup time is: `initialDelaySeconds + failureThreshold x periodSeconds` Set that to comfortably exceed the **worst** observed cold start - a node under memory pressure, a cold image cache, a slow dependency at boot - not the median. A common shape is `periodSeconds: 10, failureThreshold: 30` for a five-minute allowance. Keep `periodSeconds` small enough that a fast boot is noticed promptly, since the container is not marked ready until the next scheduled check. The payoff: liveness can now be aggressive - `periodSeconds: 10, failureThreshold: 3` gives roughly 30-second detection of a genuine hang, which would be unsafe without a startup probe. ## Choosing the handler Often the startup probe reuses the readiness endpoint, or a cheaper dedicated one. A `tcpSocket` startup probe is tempting but weak - the listener may open long before the app can serve. Prefer an HTTP endpoint that succeeds only when initialisation genuinely finished. ## Common mistakes - **Keeping a large `initialDelaySeconds` on liveness as well** - redundant and confusing; the startup probe has taken over that role. - **Making the startup budget too small**, producing a boot-time crash loop easily misdiagnosed as an application failure. The tell is a rising `restartCount` with `Startup probe failed` events and logs that always end mid-initialisation. - **Assuming the startup probe keeps running.** It does not; after the first success only liveness and readiness matter, so a container that degrades later is caught by liveness - which is why you still need one. - **Using it in place of an init container** for ordered prerequisites. A startup probe observes the app's own progress; it does not perform setup. ## Interview shape A complete answer names the old dilemma, states the three interactions (liveness disabled, readiness disabled, runs once), gives the budget formula, and closes with the practical benefit: tight liveness timings become safe.
- How long may a container take to start with periodSeconds 10 and failureThreshold 30 on its startup probe?Up to about 300 seconds, plus any initialDelaySeconds, because the kubelet retries every 10 seconds and tolerates 30 consecutive failures. Once that budget is exhausted the container is killed and restarted according to the Pod's restartPolicy.
- Does the startup probe keep running after it succeeds?No. It succeeds exactly once per container instance and is then never executed again; liveness and readiness take over. A container that becomes unhealthy later is handled by the liveness probe, which is why you still need one.
saying these in an interview costs you the question
- Thinking the startup probe runs periodically for the container's whole life
- Believing liveness still runs and can restart the container while the startup probe is pending
- Assuming a Pod can serve traffic before its startup probe succeeds
- Using a startup probe to perform initialisation instead of observing it
- Keeping a huge initialDelaySeconds on liveness after adding a startup probe