skip to content

Infrastructure automation can be written to react to change notifications, or to periodically compare the entire declared state against reality. Explain the difference between edge-triggered and level-triggered designs, and why reconciliation systems are built the level-triggered way.

level: middleimportance: nice to knowfreq 36%

answer

  1. transition versus current value
  2. function of the event or of the state
  3. a dropped event is never recovered
  4. some things go wrong with no event at all
  5. events as hints, full comparison anyway

basics

~20 s

Edge-triggered automation acts on change events; level-triggered automation repeatedly inspects the current state and closes whatever gap it finds. Reconciliation is level-triggered because a missed, duplicated or out-of-order event leaves an edge-triggered system permanently wrong, while a level-triggered one self-heals on the next pass.

solid answer

~50 s

The terms come from electronics. An edge-triggered design fires on the *transition* — something changed, here is the notification, do the corresponding work. A level-triggered design looks at the *current value* on every pass and acts on what it sees, regardless of how it got there. The difference that matters operationally is what happens when the notification stream misbehaves. Events get dropped when a delivery fails, duplicated on retry, delivered out of order, or lost entirely while the handler is down for a deploy. An edge-triggered handler that misses one is wrong until someone notices — it has no way to discover a change nobody told it about. A level-triggered loop does not care: on the next pass it compares the whole declaration against reality and repairs whatever is off, whether the cause was a missed event, a console edit, or something it did wrong itself. Events are then just an optimisation that makes the next pass happen sooner.

go deeper

for a junior

Know the plain distinction: one design reacts to notifications that something changed, the other keeps checking how things currently are. Be able to say that notifications can go missing and the checking approach recovers anyway.

for a middle

Explain that the action is a function of state rather than of the event, and give the concrete failure list — dropped, duplicated, reordered and never-emitted events — that makes the edge-triggered design decay.

for a senior

Show the hybrid in practice: events used only to schedule an earlier pass, a periodic pass underneath as the safety net, and the ability to spot edge-triggered logic from its production symptom of rare, unreproducible stuck objects.

for a principal

Set it as a design standard for the automation your organisation writes: no handler may assume it saw every event, and any cleanup cron that appears is treated as evidence the logic belongs in a reconciliation loop instead.

## Where the terms come from The vocabulary is borrowed from digital electronics. An edge-triggered input responds to a signal *changing* — the rising or falling edge. A level-triggered input responds to the signal's *current level*, continuously, for as long as it is held. Distributed-systems people adopted the pair because it names a design decision that otherwise takes a paragraph. ## The two designs **Edge-triggered automation** subscribes to a change feed and maps each notification to an action. A resource was created, so provision its companions. A tag changed, so update the inventory. The system's correctness depends on the notification stream being complete, ordered and delivered exactly once — because the action is a *function of the event*, and an event you never received produces no action at all. **Level-triggered automation** ignores how it was woken up. On each pass it reads the current declaration, observes the current world, computes the difference and acts on it. The action is a *function of the state*, not of the history. Every pass is a complete re-derivation, so the system has no memory to get wrong. ## Why level wins for infrastructure **Notification streams are unreliable in exactly the ways that matter.** Cloud change feeds are best-effort or at-least-once, not exactly-once and rarely ordered. A handler that is down during a deployment misses everything sent while it was gone, and there is no backfill. A retry delivers the same event twice. Two events arrive in the wrong order, so a handler that computed "scale to 5, then scale to 3" ends at 5. **Not every change produces an event.** Something can become wrong without anything transitioning: a certificate expires, a quota fills, a dependency is deleted in another account, a provider changes a default. Edge-triggered automation is structurally blind to those, because there is no edge to catch. A level-triggered pass sees the current condition and reacts to it. **Failure is self-healing.** If a level-triggered pass fails halfway, the next pass simply re-derives the gap and finishes. The system needs no record of "which events did I already handle" — the deduplication problem, the ordering problem and the replay problem all evaporate because there is nothing to replay. **Idempotence comes free.** A pass that runs when there is nothing to do does nothing. So it is safe to run the loop far more often than changes actually happen, which is what makes short intervals viable. ## The cost, and the hybrid everyone ends up with Level-triggered has a real price: latency and load. A pure polling loop reacts, on average, half an interval after the change, and it re-reads the whole world every pass whether or not anything moved. Over a large estate this is a lot of API calls for very few actions. So mature systems combine the two: events are used as *hints* to schedule an immediate pass, while a periodic full pass runs regardless. The event makes the reaction fast; the periodic pass makes the system correct. Crucially, the handler still does the full state comparison rather than trusting the event's payload — the event says "look now", not "do this". That distinction is the whole design, and it is the sentence to say in an interview. ## Recognising the anti-pattern Automation written as a set of handlers — one for created, one for updated, one for deleted, each performing the corresponding action — is the shape that decays. Each handler encodes an assumption about what the world looked like before the event. Bugs appear as "it works until something is missed", they are unreproducible, and the usual fix is a growing pile of reconciliation scripts run by cron to clean up after the handlers. At that point the team has rebuilt a level-triggered loop by accident, badly. Cluster controllers are the best-known systems designed level-triggered from the start, and the same reasoning applies to any infrastructure automation you write yourself. ## In the interview Give the definitions in one sentence each, then go straight to the missed-event argument — that is what the question is testing. Adding "we still use events, but only to trigger a pass sooner" shows you understand it as an engineering trade rather than a slogan.

  • If level-triggered is more robust, why use change events at all?
    For latency and cost. A pure interval loop reacts on average half an interval late and re-reads everything each pass, which is expensive over a big estate. Events let you schedule a pass immediately after something changes, so you get fast reaction with a long safety interval underneath. The rule is that the event only triggers the pass — the pass still compares full state rather than trusting the event's contents.
  • What symptom in production suggests automation was written edge-triggered?
    Correctness that decays quietly: things are right most of the time, and occasionally one object is stuck in a state nobody can explain or reproduce. It typically correlates with a deployment or an outage of the handler, and the team's workaround is a scheduled cleanup script. That script is a level-triggered loop reinvented, which is the signal to move the logic into it.

saying these in an interview costs you the question

  • Cloud change events are delivered exactly once and in order
  • Reacting to events is more efficient, so it is better
  • Every incorrect state is preceded by a change event
  • A periodic full pass is wasteful if events already fire
  • Handlers per event type are the natural way to write automation

context