How would you define and measure an index-lag SLO for a product-search index fed from a transactional database?
answer
- commit to searchable, not to written
- percentile, threshold, window
- oldest pending change, in seconds
- canary product polled from search
basics
~20 sDefine lag as the time from a change committing in the database to it appearing in search results, and target a percentile, such as 99% of changes within 30 seconds. Measure it with commit timestamps carried through the pipeline plus synthetic canary products.
solid answer
~50 sIndex lag is the delay between a change **committing** in the database and becoming **searchable**. I would write the SLO as a percentile over a window, for example 99% of changes searchable within 30 seconds over 28 days, possibly with a tighter target for stock and price. To measure it, each event carries its commit timestamp, and the indexer records `ack_time - commit_time` when the index accepts the write. The indexer also reports the age of the **oldest unprocessed** change, which catches a stalled pipeline that emits no samples at all. Because the engine's visibility delay happens after the acknowledgement, a **canary** writes a synthetic product every minute and polls search until it appears, which gives the true end-to-end figure. Alert on error-budget burn rate, not on single spikes. For stale results, the UI re-checks price and stock against the source of truth on the product page and at checkout.
go deeper
Remember that index lag means time from commit to searchable, and that search can briefly show old prices or stock.
Name the stages that add lag and explain why a percentile target and a time-based measurement beat averages and message counts.
Show the full measurement set: per-event histogram, oldest-unprocessed age, a canary through search, burn-rate alerts, and how checkout re-validates stock.
Tie the threshold to the cost of staleness per field class, and decide who owns the error budget when bulk jobs from other teams consume it.
## What index lag is **Index lag** is how far the search index trails the source of truth: the time between a product change **committing** in the transactional database and that change being **visible** to search queries. It is the sum of several stages, and each can be the bottleneck: | Stage | What adds delay | Typical cause of spikes | |---|---|---| | Capture | Outbox relay polling or commit-log reading | Relay down, log reader behind | | Queue wait | Events waiting for a consumer | Traffic burst, too few consumers | | Indexing | Transforming and writing documents | Slow enrichment calls, index throttling | | Visibility | Engine buffering new writes before exposing them | Visibility interval raised for throughput | A **service level objective (SLO)** turns that delay into a target the team can be held to. ## Writing the SLO A useful index-lag SLO has four parts: 1. **The event measured**: a committed product change becoming searchable. 2. **A percentile**, not an average, since averages hide the slow tail that users notice. Example: 99% of changes. 3. **A threshold**: within 30 seconds, as an illustrative figure chosen from what the business can tolerate. 4. **A window**: measured over a rolling 28 days, which defines the **error budget**. Here 1% of changes may exceed 30 seconds. Many teams set **field classes**: a tight target for stock, price and delisting, and a looser one for descriptions or images. The target should come from the cost of staleness, such as oversold items or wrong prices, not from what the pipeline happens to achieve today. ## Measuring it end to end No single metric covers every stage, so combine three signals: - **Per-event lag**: each event carries the database commit timestamp. When the index acknowledges the write, the indexer records `ack_time - commit_time` in a histogram. This covers capture, queue and indexing, but not the visibility delay. - **Oldest-unprocessed age**: the indexer reports how long the oldest pending change has waited. If the pipeline stalls completely, no events are acknowledged and the histogram goes quiet, so this is the signal that catches outages. - **Synthetic canary**: a job writes a dedicated test product with a unique token every minute or so, then polls search until the token is returned. The elapsed time is the true end-to-end lag, including the engine's visibility delay. Canary products must be excluded from real results and analytics. A message backlog count is a weak substitute: 10,000 pending events might be two seconds or two hours of delay depending on throughput. Report lag in **time**, not in messages. The commit timestamp and the acknowledgement time come from different machines, so clock skew adds error. It is usually small next to a 30-second target, but worth knowing when the target is tight. ## Alerting on the SLO - Alert on **burn rate**, meaning how fast the error budget is being used, over both a short and a long window. That pages on real incidents without paging for every brief spike. - Alert separately on oldest-unprocessed age crossing the threshold, since a stalled pipeline may consume budget before the histograms show it. - Break lag down by stage in dashboards, so the on-call engineer can see which stage is behind. ## What the product does about stale results An SLO limits staleness; it does not remove it. The UI and the flows around search must assume results can be a little old: 1. **Re-validate at decision points**: the product page and checkout read price and stock from the source of truth, so search can be stale without selling something that is gone. 2. **Read your own writes**: after a seller edits a listing, their own dashboard reads from the database or overlays the pending change, so they do not think the edit was lost. 3. **Degrade visibly during a breach**: hide the stock badge, or add an availability check when results are rendered, instead of showing confident but wrong data. 4. **Filter at query time** for critical states such as delisted products, where a short list of recently removed IDs can be applied to results until the index catches up. The measurement tells you when the index is behind; these patterns decide how much that matters to a shopper.
- Why is per-event lag recorded at acknowledgement not enough on its own?It stops at the moment the index accepts the write, so it misses the engine's visibility delay before the document is searchable. It also goes silent during a full stall, because no events are acknowledged and no samples are recorded. The oldest-unprocessed age covers the stall, and a synthetic canary polled through search covers the visibility delay.
- How should a seller who just edited a listing avoid seeing the old version in search?Apply read-your-writes for that seller: their own listing views read from the source of truth, or the client overlays the edit it just made until the index reports a version at least as new. Other shoppers can tolerate the few seconds of lag; the editor is the one who notices and reports the change as lost.
- The SLO is breached during a nightly bulk price update; what would you change?Separate the traffic: send bulk updates through their own lower-priority stream or consumers, so they cannot delay individual stock and price changes on the main path. Also scale consumers for the known window, and consider per-field SLO classes so that a bulk description rewrite is not measured against the stock target.
saying these in an interview costs you the question
- A message backlog count alone tells you how stale the index is.
- Average lag is a good enough SLO for index freshness.
- Measuring until the index acknowledges the write covers end-to-end lag.
- With a good SLO, checkout can trust search results for stock.
- A lag histogram going quiet means the pipeline is healthy.