How do you choose periodSeconds, timeoutSeconds, failureThreshold and initialDelaySeconds for a Kubernetes probe? Work through what those numbers mean for detection time and for false restarts under load.
answer
- detection = initialDelay + period x failureThreshold
- timeoutSeconds default 1 - the classic footgun
- liveness forgiving, readiness twitchy
- successThreshold must be 1 for liveness/startup
- probe load = replicas / period
basics
~20 sWorst-case reaction time is roughly initialDelaySeconds + periodSeconds x failureThreshold, plus timeoutSeconds. Defaults are period 10s, timeout 1s, failureThreshold 3. Keep liveness generous and cheap so latency spikes do not restart healthy Pods; keep readiness tighter, since removing traffic is reversible.
solid answer
~60 sThe four fields combine into one number: **time to react = `initialDelaySeconds` + `failureThreshold` x `periodSeconds`**, plus up to `timeoutSeconds` per attempt. Defaults: `periodSeconds: 10`, `timeoutSeconds: 1`, `failureThreshold: 3`, `successThreshold: 1` (and 1 is the only legal value for liveness and startup). How I set them: - **Liveness: forgiving.** It is destructive, so it should fire only on a real hang - period 10, failureThreshold 3-5, timeout 3-5 seconds. The 1-second default timeout is the classic footgun: a GC pause or a busy node makes a healthy JVM miss it, and three misses in a row restart it. - **Readiness: tighter.** Removing traffic is reversible, so a shorter period and threshold (5s / 2) let a saturated instance shed load quickly and rejoin quickly. - **A startup probe instead of a big `initialDelaySeconds`**, so boot tolerance does not blunt steady-state detection. And probes must be cheap: cost scales with `replicas / periodSeconds`, and a health endpoint that queues behind the same thread pool as real traffic turns overload into a restart storm.
code
yaml · 15 lineslivenessProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: 8080
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2
successThreshold: 1go deeper
Know the field names, the defaults, and that failureThreshold consecutive failures are needed before Kubernetes acts.
Compute detection time from period and threshold, and explain why the 1-second default timeout causes false failures.
Tune asymmetrically - forgiving liveness, reactive readiness - justify it by the destructiveness of a restart, and recognise probe-induced restart storms in events and metrics.
Own the fleet defaults and the health-endpoint contract, budget aggregate probe load, and weigh detection latency against the blast radius of unnecessary restarts across the estate.
## The fields - `initialDelaySeconds` (default 0) - wait this long after the container starts before the first check. - `periodSeconds` (default 10) - interval between checks. - `timeoutSeconds` (default **1**) - how long one check may take before it counts as a failure. - `failureThreshold` (default 3) - consecutive failures needed to act. - `successThreshold` (default 1) - consecutive successes needed to be healthy again; must be 1 for liveness and startup probes. The arithmetic every candidate should do out loud: with `periodSeconds: 10` and `failureThreshold: 3`, a hung process is restarted **20-30 seconds** after it hangs - the first failing check can land anywhere within the period, and each attempt can add up to `timeoutSeconds`. Halving the period halves detection time and doubles probe load. ## The 1-second timeout trap `timeoutSeconds: 1` is the most damaging default. A healthy service can exceed it during a stop-the-world GC pause, a cold JIT phase, a noisy neighbour on the node, or a burst of real traffic. Three such misses and the kubelet restarts a perfectly healthy container - dropping its warm caches and connections, pushing its traffic onto the remaining replicas, and making *them* more likely to miss their probes. That is the anatomy of a probe-induced cascading restart, and it always appears at peak load, the worst possible moment. Mitigations: raise `timeoutSeconds` to a few seconds for liveness, keep the health endpoint off the request-serving thread pool (or ahead of the queue), and never let the probe do real work. ## Liveness versus readiness tuning Because the consequences differ, the numbers should differ: - **Liveness** is destructive and irreversible: bias towards *false negatives*. Longer period, higher threshold, generous timeout. If unsure, prefer detecting a hang in 60 seconds over restarting a healthy Pod once a week. Many teams run no liveness probe at all for services that crash cleanly on fatal errors. - **Readiness** is cheap and reversible: bias towards *fast reaction*. A short period (2-5s) and low threshold (2) let an overloaded instance shed traffic immediately, and a `successThreshold` above 1 prevents flapping straight back into rotation. ## Startup probe versus initialDelaySeconds A large `initialDelaySeconds` on liveness to survive a slow boot also blinds you for that long after *every* restart, and it does nothing for the variance of a genuinely slow start. A `startupProbe` sized to the worst-case cold start is the right tool; then `initialDelaySeconds` on liveness can be 0. ## Probe load Total probe rate is `replicas x containers / periodSeconds`. At 500 replicas and a 1-second period that is 500 checks per second, plus whatever each check touches. If the endpoint queries a cache, a dependency or the disk, probes become a meaningful share of the service's own load - and an exec probe additionally forks a process each time. Cheap, constant-time health endpoints are not a nicety. ## A worked default For a typical HTTP service: - `startupProbe`: httpGet, period 10, failureThreshold 30 - five minutes of boot allowance. - `livenessProbe`: httpGet `/healthz` with local-only checks, period 10, timeout 3, failureThreshold 3 - roughly 30-second detection. - `readinessProbe`: httpGet `/readyz` including required dependencies and warm-up, period 5, timeout 2, failureThreshold 2 - about 10 seconds to leave rotation. Then validate: force a hang in staging and measure real detection time; run a load test and confirm no probe failures at peak; and after any incident, check whether probes helped or amplified it. ## Signals that the tuning is wrong Restarts clustering at peak traffic, `Unhealthy` events mentioning timeouts, restart counts rising across *all* replicas simultaneously, or Pods flapping in and out of endpoints. All four point at probe configuration rather than at the application.
- With periodSeconds 10 and failureThreshold 3, how long after a process hangs is the container restarted?Roughly 20 to 30 seconds. The hang may occur just after a successful check, so the first failure can be up to one period away, and two more failed periods are needed after that; each attempt can additionally take up to timeoutSeconds.
- Your Pods restart only during peak traffic. How do you tell whether the app or the probe is at fault?Look at the Unhealthy events and the previous termination reason: probe-induced restarts show liveness timeouts across many replicas at once with no application error in the previous logs. Confirm by checking whether the health endpoint shares the saturated thread pool, then raise timeoutSeconds and failureThreshold and re-test under load.
Tuning liveness is like setting a smoke alarm's sensitivity: too twitchy and it goes off every time someone makes toast, evacuating a perfectly good building at exactly the busiest moment.
saying these in an interview costs you the question
- Leaving timeoutSeconds at the 1-second default for a JVM or any GC-pausing runtime
- Setting a very short period and low threshold on liveness to 'detect faster'
- Using a big initialDelaySeconds on liveness instead of a startup probe
- Setting successThreshold above 1 on a liveness probe, which is invalid
- Ignoring that probe cost multiplies by replica count