In Zabbix, what does a trigger expression evaluate, and what does a recovery expression add?
answer
- a condition over stored values
- functions read a window of history
- min over five minutes, not last
- a second expression closes the problem
- a deadband stops the flapping
basics
~20 sA Zabbix trigger is a boolean expression over an item's stored history, built from functions like min or avg over a window rather than a threshold on the latest value. A separate recovery expression clears the problem on a looser condition.
solid answer
~50 sA Zabbix **trigger** is a named boolean expression that reads an item's stored history through trigger functions — `min`, `max`, `avg`, `count`, `change`, `nodata` — over a time period or a count of values. `min(/ferry-booking-web-03/system.cpu.util,5m)>90` opens only if utilisation stayed above 90% for the whole five minutes; `last()` is the one function that behaves like a naive threshold on the newest value. By default the problem closes when that expression stops being true. Setting the trigger's OK event generation to **Recovery expression** lets you supply a second, independent condition such as `max(...,10m)<75`. The gap between the two bounds is a **deadband**, and it is the fix for a metric parked near its limit: without one, values oscillating either side of 90 open and close the problem on every crossing, producing dozens of fragmented problems instead of a single readable incident.
code
text · 2 linesProblem: min(/ferry-booking-web-03/system.cpu.util,5m)>90
Recovery: max(/ferry-booking-web-03/system.cpu.util,10m)<75go deeper
Recall that a Zabbix trigger is a named expression with two states, PROBLEM and OK, and that it reads values an item already collected rather than holding data of its own.
Be able to read an expression such as min over a five-minute window and say what it means, name the common trigger functions including nodata, and explain what switching OK event generation to a recovery expression changes.
Show you have fixed a flapping trigger in production: why raising the threshold fails, how a deadband between the problem and recovery bounds fixes it, and what fragmented problem history did to the on-call record.
Own the trigger conventions for the whole estate: which functions are permitted, how deadbands are chosen rather than guessed, and how trigger dependencies stop one root fault generating a hundred derived problems.
## A trigger is an expression, not a threshold A Zabbix **item** collects values; it never decides anything. The decision lives in a **trigger**: a named boolean expression attached to one or more items, with two states, `OK` and `PROBLEM`. Zabbix re-evaluates a trigger when a new value arrives for an item it references, and expressions built on time-based functions are re-evaluated periodically as well, so they can fire when nothing arrives at all. The expression is not "value greater than number". It is built from **trigger functions**, each of which reads a slice of the item's stored history: ``` min(/ferry-booking-web-03/system.cpu.util,5m)>90 ``` Read it left to right: take the item `system.cpu.util` on host `ferry-booking-web-03`, look at every value collected in the last five minutes, take the smallest, and compare it with 90. That trigger opens only if utilisation stayed above 90% for the *whole* window, so a single spike cannot raise it. Swap `min` for `avg` and you have made a weaker statement; swap it for `last` and you are back to a bare threshold on the newest value. The common functions are worth knowing by name: - `last()` — the most recent value; the only one that behaves like a naive threshold. - `min()`, `max()`, `avg()` over a period — the sustained-condition family. - `count()` — how many values in the period matched a condition. - `change()` — the difference between the last two values. - `nodata()` — true when *no* value has arrived for the period. This is how a silent host or a stalled collection path is detected, and nothing based on the latest value can express it, because there is no latest value. Functions take either a time period (`5m`, `1h`) or a count of values (`#3` for the last three), and one expression may reference items on several hosts, which is how "the queue is deep *and* the database is lagging" becomes a single problem rather than two correlated ones. ## What a recovery expression buys you By default a trigger returns to `OK` when its problem expression stops being true — the setting Zabbix calls **OK event generation: Expression**. Change it to **Recovery expression** and you supply a second, independent expression that must become true before the problem closes. The reason is **hysteresis**. Consider a metric parked near its limit: CPU on the ferry-timetable booking API sitting between 89% and 92% through the evening sailing rush. With one expression at `>90` the trigger opens and closes on every crossing — forty-one state changes in an hour, forty-one problems, and a problem history nobody can read. Raising the threshold does not help; it just moves the line the metric oscillates around. A recovery expression separates the two edges: | | Single expression | Problem plus recovery expression | |---|---|---| | Opens when | `min(...,5m)>90` | `min(...,5m)>90` | | Closes when | the same expression is false | `max(...,10m)<75` | | Behaviour near the limit | flaps on every crossing | one problem, held until clearly recovered | | Duration reported | fragmented into slivers | the real duration of the incident | | Failure mode | alert fatigue | a problem open slightly longer than the fault | The gap between 90 and 75 is a **deadband**: while the metric is inside it, the trigger keeps whatever state it already has. Note that the recovery expression is deliberately *not* the negation of the problem expression — a candidate who writes `max(...,5m)<=90` as the recovery has re-created exactly the flapping they set out to remove. The same mechanism covers cases where recovery is a genuinely different measurement rather than a looser bound: open on `nodata()` for an agent availability item and close on the first successful value; open when a queue depth exceeds a limit and close only after it has been drained for a sustained window. ## Getting the shape right A handful of choices decide whether a trigger is usable in production: 1. **Choose the function before the number.** Most bad triggers are `last()` with a carefully tuned threshold where `min()` or `avg()` over a window was wanted. 2. **Set severity honestly.** Zabbix offers Not classified, Information, Warning, Average, High and Disaster. Severity is a property of the trigger, and everything downstream filters on it. 3. **Use trigger dependencies.** When a host is unreachable, its twenty per-service triggers are noise; making them depend on the host-unavailable trigger suppresses the derived problems while the root problem is open. 4. **Reference items, never duplicate them.** A trigger stores nothing of its own, so an expression whose window is longer than the referenced item's history retention silently evaluates over less data than you think. 5. **Watch the evaluation cost.** A long window over a high-frequency item is recalculated on every arriving value; hundreds of those across a fleet is real server work, not a rounding error.
- What does the nodata() function let a Zabbix trigger detect that no threshold can?Absence. Every other trigger function needs values to compute over, so a host that stops reporting entirely produces no evaluation and therefore no problem. nodata() is true when nothing has arrived for the given period, which is how a dead agent, a stopped service or a silently broken collection path gets noticed at all.
- Why does raising the threshold not fix a flapping Zabbix trigger?Because flapping is caused by the metric crossing a single line repeatedly, not by where that line sits. Move it to 95 and a metric that now oscillates around 95 flaps just as badly. The fix is two lines: the problem expression at one bound and a recovery expression at a lower one, with a deadband in between where the trigger simply holds its current state.
- Can one Zabbix trigger reference items on more than one host?Yes. Each function names its item with a host and key path, so a single expression can combine, for example, queue depth on one host with replication lag on another. That is useful for expressing one real fault as one event rather than two correlated ones, but it couples the trigger to data arriving from both hosts.
A recovery expression is the deadband on a thermostat: it turns the heating on at eighteen degrees and off at twenty-one, rather than switching every time the room drifts across a single number.
saying these in an interview costs you the question
- Says a Zabbix trigger fires when the latest value crosses a number
- Fixes a flapping trigger by raising the threshold instead of adding hysteresis
- Writes the recovery expression as the plain negation of the problem expression
- Thinks a trigger stores its own data rather than reading item history
- Believes a trigger belongs to a single item and cannot span hosts