skip to content

In Zabbix, where does the server bottleneck as hosts are added, and what does a proxy change?

level: seniorimportance: nice to knowfreq 33%

answer

  1. values per second, not host count
  2. a fixed pool of collecting processes
  3. the queue is the early symptom
  4. collection moves out, storage does not
  5. buffers locally through a link outage

basics

~20 s

A Zabbix server is sized by new values per second, not host count. Collection saturates first: pollers are a fixed pool, so checks are delayed and queue up. A proxy moves collection off the server and reaches networks it cannot address.

solid answer

~50 s

A Zabbix server is sized by **new values per second** — item count times frequency — not by host count. For each value it polls or receives, preprocesses, writes into a history cache flushed by history syncers, and recalculates every trigger referencing that item. **Collection saturates first.** Pollers are a fixed pool set by `StartPollers`; once demand exceeds it, checks are delayed rather than dropped and the backlog shows up as a growing queue. Watch `zabbix[queue,10m]` alongside `zabbix[process,poller,avg,busy]`. Unreachable hosts make it far worse, since each blocks a poller for its full timeout. The database write path is the next and harder ceiling. A **proxy** collects and preprocesses for an assigned set of hosts and forwards the values, so it removes poller load and reaches networks the server cannot address, buffering through link outages in its own database. It does not reduce database load, and it does not evaluate triggers.

code

text · 2 lines
text
zabbix[queue,10m]
zabbix[process,poller,avg,busy]

go deeper

for a junior

Recall that a Zabbix server collects, stores and evaluates, and that a proxy is an extra component doing the collecting for a group of hosts and forwarding the values on to the server.

for a middle

Explain why values per second rather than host count sizes the server, what a poller does for a passive check, and what the queue is actually measuring when it starts to grow.

for a senior

Show you have diagnosed it: separating poller saturation from a database write ceiling, using the internal queue and process-busy items as evidence, and knowing that unreachable hosts starve pollers through their timeouts.

for a principal

Own the topology. Decide how many proxies and where, accept that they move collection but not storage, and be explicit about what deferring trigger evaluation to the centre means for detecting a remote site's outage.

## What the server actually does for each host A Zabbix server's workload is not measured in hosts. It is measured in **new values per second**, which is the product of item count and collection frequency. A host with nine items on a one-minute interval costs a fraction of what a host with sixty items on a ten-second interval costs, and two estates with identical host counts can differ by an order of magnitude. For every value, the server does roughly this: 1. **Collect it.** A passive item occupies a **poller** process for the whole round trip: connect to the agent, send the key, wait, read, close. If the host is slow or unreachable that poller is blocked until the timeout expires. An active item costs far less — the agent sends batches to a trapper process, which only has to accept them. 2. **Preprocess it.** Change-per-second calculations, JSONPath extraction, regular expressions, unit conversion and throttling all run in preprocessing workers before the value is stored. 3. **Write it.** The value lands in an in-memory history cache and is flushed to the database by history syncer processes, which also maintain trends. 4. **Decide on it.** Every trigger referencing that item is recalculated, reading history through the value cache. Only then can a problem be raised. 5. **Keep configuration in reach.** Hosts, items, triggers and macros are held in a configuration cache refreshed periodically, and that cache grows with the estate. ## Where it breaks first Bring up a new region for the ferry-timetable booking platform — 1,847 hosts and 41,600 items appearing at once with nothing tuned — and the failure is almost always the same, in almost always this order. **Poller saturation comes first.** Pollers are a fixed pool sized by `StartPollers`. When demand exceeds the pool, checks are not dropped, they are *delayed*, and the delay is visible as the **queue**: items whose collection time has passed and which have not been collected yet. A queue that grows and never drains is the clearest signal that collection capacity is short. Zabbix monitors itself here through internal items — `zabbix[queue,10m]` counts items overdue by more than ten minutes, and `zabbix[process,poller,avg,busy]` reports how much of its time the poller pool spends busy. A pool sitting at 96% busy has no headroom whatever the queue reads at that instant. The classic aggravator is unreachable hosts. One dead network segment means every passive check into it burns a poller for the full timeout, so a handful of dead hosts can starve collection for healthy ones. Zabbix mitigates this with a separate pool for hosts it has already marked unreachable, which is a setting worth knowing exists before you need it. **The database is the second and harder ceiling.** History and trend writes are relentless, and no amount of extra pollers helps if the syncers cannot flush them. The symptom is different: the history cache's free percentage falls, and if it reaches zero the server stalls at the front door. Behind that sits the housekeeper deleting expired history, which on a large installation is itself significant load — one reason large estates move history to a time-series-optimised backend and lean on partition drops rather than row-by-row deletion. **Trigger evaluation is third**, and usually only when expressions are careless: long windows over high-frequency items, recalculated on every arriving value. ## What a proxy changes A **Zabbix proxy** is a separate process with its own database that collects on behalf of a set of hosts assigned to it and forwards the values to the server. It comes in two flavours: an **active proxy** connects out to the server, and a **passive proxy** waits for the server to connect to it — the same direction question as an agent's checks, one level up. | Moves to the proxy | Stays on the server | |---|---| | Polling and trapping for its hosts | Trigger evaluation and problem events | | Item preprocessing | History and trend storage | | Buffering while the link is down | Configuration and the frontend | | Network reach into that segment | Everything a human interacts with | There are two distinct wins, and interviews test whether you can separate them: - **Load.** Polling and preprocessing leave the server, so the poller ceiling stops being the constraint and collection scales by adding proxies. What does *not* change is the database: the same values still arrive and are still written to the server's history. A proxy is not a data-reduction device, and expecting it to shrink storage is the most common misconception on this subject. - **Reach.** A proxy placed inside a remote site, a DMZ or a customer network monitors hosts the server cannot address at all, over a single connection instead of a firewall rule per host. Because it keeps its own database, it continues collecting and buffering through a link outage and back-fills when the link returns. The limit is worth stating plainly: **a proxy does not evaluate triggers.** During a link outage the remote site keeps collecting, but nothing is decided until the values reach the server, so a site's outage is detected centrally and late rather than locally and immediately. Sizing follows the same logic as the server itself — count the values per second a proxy will carry, not the hosts assigned to it.

  • Does adding a Zabbix proxy reduce load on the server's database?
    Barely. The proxy takes over polling and preprocessing, but every value it collects is still forwarded to the server and written into the same history and trend storage. A proxy scales collection and extends network reach; it is not a data-reduction device. To cut database load you have to collect less, sample less often, or store history differently.
  • How do you tell a saturated poller pool apart from a database write ceiling in Zabbix?
    They show up in different places. Poller saturation appears as a high busy percentage on the poller processes together with a growing queue of overdue items. A write ceiling appears as the history cache's free space falling while collection itself is keeping pace. Adding pollers helps the first case and actively makes the second one worse.

A proxy is a branch office: it does the collecting and keeps the paperwork safe when the road to head office is washed out, but head office still makes every decision.

saying these in an interview costs you the question

  • Thinks a Zabbix proxy reduces the volume written to the server's database
  • Says a Zabbix proxy evaluates triggers for the hosts it collects from
  • Treats a growing queue as a database problem without checking poller utilisation
  • Adds pollers without checking whether the write path can absorb the values
  • Sizes a Zabbix server by host count while ignoring item count and intervals