A device-telemetry ingester's spec carries a bad setting, so every replacement comes up faulty - what does self-healing actually restore?
answer
- one objective: the count
- the spec is the premise
- faults are copied, not caught
- full count, zero serving
- fix the spec, not the copies
basics
~20 sOnly the count, measured against the spec as written. The loop starts copies until the declared number exists; it never edits the spec, so a fault that lives in the spec is reproduced identically in every replacement, and a workload can sit at full count serving nothing.
solid answer
~50 sSelf-healing is convergence on one number: how many copies match the declared workload against how many should. The loop reads the spec and starts instances of it - it has no notion of whether the spec is right. A wrong endpoint, an unreachable image reference, a missing credential or a bad flag is baked into every replacement exactly as it was into the original, so the platform either churns forever or parks at the declared count with every copy useless. This is why `count == declared` is not a health signal. The signals that catch it are the ones measured on serving rather than on existence: how many copies report themselves able to serve, how long they survive, and what the workload's own output shows. A broken spec is fixed by changing the spec; no amount of looping converges on correctness.
code
yaml · 12 lines# declared by the team - the only input the loop reads
spec:
replicas: 5
imageDigest: "<content digest of the built image>"
settings:
ingestEndpoint: "telemetry-intake.internal:9443" # wrong port
# read back from the platform - what it measured, not what it repaired
status:
observedReplicas: 5 # the loop has converged
readyReplicas: 0 # nothing is serving
lastReplacementAt: "03:14"go deeper
Hold onto one sentence: the platform restores how many copies exist, using the spec exactly as written. If the spec is wrong, every new copy is wrong the same way.
Explain why the loop cannot evaluate its own input, and name the three shapes the failure takes: nothing starts, copies start and die, or copies live and serve nothing.
Show which signals you would watch instead of the count - serving copies, instance age, the workload's own throughput - and say how you would distinguish a bad spec from a failing dependency.
Argue for where correctness checks belong given that the loop cannot provide them: what the update path should verify before advancing, and what the team owes in signals for workloads whose count is meaningless.
## What the loop is optimising The replica loop has exactly one objective function: make the number of copies matching this workload's selector equal the number declared in its spec. Everything it does follows from that. It does not evaluate the spec, does not compare this version of the spec to the last one that worked, and has no concept of "this configuration is wrong". The spec is the premise of its reasoning, not an input to be checked. So when the spec itself carries the fault, the loop faithfully reproduces the fault. Three shapes of this are common: - **The replacement never starts.** The image reference cannot be resolved or fetched, so nothing runs. The count stays short and the loop keeps trying. - **The replacement starts and dies.** A missing credential or an unreachable dependency makes each copy exit shortly after start. The count oscillates, and the platform's in-place restart behaviour with its escalating delay takes over on whichever host it landed. - **The replacement starts, stays up, and does nothing useful.** The worst case: the ingester points at the wrong upstream endpoint, so the processes are alive, the declared count is satisfied, and no telemetry is being ingested at all. That third case is the one interviews are really probing, because the platform reports a fully converged workload while the service is down. ## Count is not health | What the loop measures | What it cannot tell you | |---|---| | How many copies exist that match the selector | Whether any of them can serve a request | | That the declared number is now satisfied | Whether the declared number is the right number | | That a new copy was successfully started | Whether the spec it was started from is correct | | That the spec was applied as written | Whether the settings inside it point anywhere real | The useful counter-signals are the ones measured on behaviour rather than existence: 1. **Copies reporting themselves able to serve**, as distinct from copies that exist. A workload at full count with zero serving copies is the exact fingerprint of a bad spec. 2. **Instance age.** If every copy is under two minutes old, replacements are churning even though the count keeps returning to target. 3. **The workload's own output.** For an ingester, readings accepted per minute is the only thing that actually says it works; process count says nothing about it. ## Why the loop does not, and should not, try harder It is tempting to want the platform to notice that the last five replacements all failed and roll something back on its own. The reason it does not is that the declared spec is the source of truth by construction. A loop that edited it would be deciding that its own guess beats what was declared, and the next pass would then be comparing against a state nobody wrote down. The model only works because the spec is the fixed point. What platforms do offer instead sits one layer up and is deliberately separate: a stepwise replacement under a new spec can watch whether new copies come up healthy and stop advancing when they do not, leaving the old copies in place. That is a property of how an update proceeds, not of self-healing - and it only helps when there was a previously working version to hold onto. For a workload whose first-ever spec is wrong, there is nothing to hold. ## The failures no loop can heal Stated plainly, automatic replacement recovers exactly one class of problem: **an instance that is gone when the spec says it should exist**. Outside that class it does nothing useful, and sometimes it obscures: - A fault in the declared spec - wrong image reference, wrong setting, missing configuration - is copied into every replacement. - A fault in something the workload depends on - an upstream that is refusing connections, a credential that expired - is unaffected by having more copies. - A fault in the application's own logic reappears in the replacement the moment it does the same work. - A capacity problem is not a loss problem: if the copies exist but cannot keep up, the loop is satisfied and the queue keeps growing. The discipline that follows is simple. Read the count as *the loop's own status*, not as the workload's health. Treat repeated replacement as a defect report about the spec or its dependencies, not as the platform coping. And when a spec is wrong, fix the spec - it is the only input the loop reads, so it is the only place a fix can take effect.
- Why does the platform not roll the workload back to the last spec that worked?Because the declared spec is the fixed point the whole model rests on. A loop that rewrote it would be preferring its own inference to what was declared, and subsequent passes would compare against a state nobody wrote down. Holding a previous version back is a property of how an update advances, and it is a separate mechanism.
- Which signal would have caught the wrong endpoint fastest?The count of copies reporting themselves able to serve, watched against the declared count. A workload sitting at five existing and zero serving is the fingerprint of a bad spec. The workload's own throughput - readings accepted per minute - catches it too, and is what proves the fix worked.
- Does self-healing help at all when the ingester's upstream dependency is down?No. Replacement only addresses copies that are missing, and the dependency is unaffected by how many copies exist. If each copy exits on the failed dependency, replacement adds start-up load to a system that is already failing, which is why an escalating delay between attempts exists.
saying these in an interview costs you the question
- Says self-healing means the platform repairs broken workloads
- Treats the declared count being met as proof of serving
- Expects the loop to correct or roll back a bad spec on its own
- Assumes churning replacements will eventually succeed by themselves
- Thinks a workload at full count cannot be an outage