How does an exploit that crashes the service one attempt in five limit a worm's spread?
answer
- no operator is watching to retry
- the pool is shared between all infected nodes
- only the first attempt against a host counts
- candidates reached times success rate
- below one, each generation shrinks
basics
~20 sEvery failure subtracts a host from the pool that all infected nodes share. A host killed on the first attempt against it is never recruited by anyone, so per-attempt success caps total reach and cuts the number of hosts each infected node can pass the code to.
solid answer
~50 sReliability is a precondition of self-spread, not polish. Self-spreading code has no operator to retry, and on many services a failed attempt kills the listener - so the failure does not merely waste a try, it permanently removes a candidate from the pool every infected node is drawing from. A host survives to be infected only if the first attempt made against it succeeds, so with 80 percent per-attempt success roughly a fifth of the reachable set is destroyed rather than recruited, and those dead hosts never carry the code onward either. What decides whether spread continues is the branching factor: candidates each infected host can reach, times the success rate. On a dense flat segment that product stays comfortably above one and the crash rate only caps coverage; where each host finds only a handful of live candidates, it drops below one and propagation dies out.
go deeper
Know that self-spreading code has no operator to retry a failed attempt, and that a failure which crashes the target service removes that machine from the set of hosts the code could ever reach.
Explain the shared-pool argument and the first-attempt rule, and give the branching-factor product - candidates reachable per host times success rate - as the thing that decides whether spread continues or dies out.
Bring the caveats unprompted: automatic restarts partly heal the pool, credential-based spreading has a non-destructive failure mode, and the crash rate matters most where each host reaches few live candidates.
Be able to argue why a funded operator may reject self-spread precisely because the reliability term is uncontrollable across a heterogeneous estate, and what they buy instead with operator hours.
## Why a hands-on operator tolerates unreliability and self-spreading code cannot An operator working host by host treats a technique that works four times in five as perfectly usable. They watch the attempt, and when it fails they retry, pick a different route, or leave that host alone. The unreliability costs them a few minutes. Self-spreading code has none of that. There is no one to notice the failure, no fallback route unless it was engineered in advance, and - the part candidates usually miss - the failure is frequently **destructive**. Techniques that abuse a memory-handling defect in a listening service typically leave the service dead when the attempt lands wrong: the process crashes, and until something restarts it the host stops listening. The attempt did not just fail; it removed the target. ## Poisoning a shared pool The crucial structural point is that every infected host is drawing candidates from the **same** finite set - on a flat unmanaged segment, the couple of hundred devices in the broadcast domain. When one infected device crashes a printer's listening service, that printer is no longer available to any of the other infected devices either. The loss is not local to the node that made the attempt. That gives a simple and slightly counter-intuitive result. A host is recruited only if the **first** attempt made against it succeeds; a first attempt that fails kills the service, and later attempts by other nodes find nothing to talk to. Retrying does not help. So with per-attempt success `p`, roughly `p` of the reachable set ever runs the code and roughly `1 - p` is destroyed. At `p = 0.8`, about a fifth of the segment is turned into wreckage instead of into forwarders. ## Whether it merely caps reach or actually stalls What decides between "caps coverage" and "stops" is the **branching factor**: how many live candidates one infected host can reach, multiplied by the success rate. - **Dense range.** Each infected device can reach a couple of hundred live neighbours. Even at `p = 0.8` the branching factor is enormous, so the crash rate costs reach - the fifth that died - and some speed, but the wave still crosses the segment. - **Sparse range.** Each infected host finds only a handful of live candidates, and some of those have already been taken by another node. Now the product `p x k` is close to one, and a crash rate that removes a fifth of the pool can push it under one. Below one, each generation is smaller than the last and propagation dies out on its own. This is why unreliability is a governing term rather than a detail: it does not scale the spread down smoothly, it moves it across a threshold. ## Two nuances worth raising unprompted **Failure is not always permanent.** Embedded and appliance-class devices - the ones filling a flat unmanaged segment - often run their listeners under a supervisor or a hardware watchdog that restarts them, sometimes within seconds. There the crash costs a window rather than a host, and the pool partially heals. Conversely a device that needs a physical power cycle, on a segment nothing reaches to restart it remotely, is gone for the entire campaign. **A crash may not be the failure mode at all.** Where the technique is a valid login with a carried credential, a wrong guess costs nothing but an authentication failure and possibly a lockout counter on that account. The target keeps listening, so the pool is not poisoned - one reason credential-based spreaders can be far more relentless than exploit-based ones despite being technically unremarkable. ## How to answer this in the room The strong answer names three things in order: failures subtract from a **shared** pool rather than from one node's list; recruitment depends on the **first** attempt against a host, so retries do not recover it; and whether that ends the spread depends on the **branching factor**, not on the crash rate alone. Then add the honest caveat about restarts. What weak answers do instead is treat reliability as a quality issue - "they would tidy that up before release" - which misses that the failure rate is arithmetic acting directly on the propagation, not an engineering blemish.
- Why do retries not recover a host that a failed attempt crashed?Because the crash removes the listener, not the attacker's chance. A later attempt from a different infected node finds nothing accepting connections on that port, so it fails for a new reason. The host is recruited only if the first attempt made against it - by anyone - succeeded, which is why per-attempt success rate caps total reach.
- How does the picture change when the spreader authenticates instead of exploiting?A failed authentication normally leaves the service running, so the pool is not poisoned and repeated attempts stay possible. The cost moves elsewhere: account lockout thresholds, and the fact that the credential either works estate-wide or not at all. Reliability stops being the limiting term and reachability of the service becomes the binding one.
- Where does a device watchdog change the calculation?It converts a permanent loss into a temporary one. Appliance-class devices often restart a crashed listener automatically within seconds, so the candidate returns to the pool and a later attempt can succeed. The spread is slowed rather than capped. A device that only comes back after a manual power cycle is a genuine permanent subtraction.
saying these in an interview costs you the question
- Treats reliability as polish rather than a precondition
- Assumes a later retry can recover a crashed target
- Thinks each infected host has its own private target pool
- Says a 20 percent failure rate simply makes spread 20 percent slower
- Ignores that crashed hosts also stop being forwarders