An availability SLI for a web API can be computed from load-balancer logs, from counters inside the application, from client-side telemetry, or from synthetic probes. What does each vantage point miss, and how do you choose between them?
answer
- each point is blind upstream of itself
- a dead process reports no failures
- clients that cannot reach you cannot report
- probes cover DNS, but at low resolution
- numerator and denominator from one place
basics
~20 sEach vantage point sees a different slice of the request path. Application counters miss everything that never reached the process, load-balancer logs miss DNS and network failures, client telemetry only reports from clients healthy enough to report, and probes measure a synthetic journey at low resolution.
solid answer
~50 sThe rule is to measure as close to the user as you can still hold yourself accountable for, which in practice makes load-balancer or ingress logs the default: they are outside the application, so they still count requests when the process is crash-looping, and they capture connection resets and capacity rejections the app never sees. In-process counters are the trap — a dead instance emits nothing, so the ratio reads a perfect 100% during the worst outage. Client-side telemetry is the truest picture of experience but suffers survivorship bias, since a client that cannot reach you often cannot report either, and it folds in the user's own network, which you cannot fix. Synthetic probes are the complement: they cover DNS, TLS and the edge, and they keep producing signal on a low-traffic service, but a one-per-minute probe cannot resolve a 0.5% error rate. I pick one primary vantage point for the committed indicator and use the others as corroboration rather than mixing them into a single ratio.
go deeper
Know that where you measure changes what you see, and that a metric emitted by the application itself disappears when the application does. Be able to name the load balancer as a vantage point outside the process.
Compare the four vantage points concretely and name one failure class each one misses — the crash-looping process for in-process counters, DNS and edge for the load balancer, the user's own network for client telemetry, and resolution for probes.
Demonstrate the operational judgment: pick one defensible primary source, keep the others as corroborating signals, refuse to blend them into one ratio, and explain what a divergence between client-side and ingress numbers is telling you.
Own the accountability boundary — how far toward the user the organisation is willing to commit, given that parts of the path are outside its control. Be able to argue what it would cost to move the committed measurement point outward and whether that buys anything a second signal would not.
## The path a request actually takes Between a user and your code sit several hops, each of which can fail: DNS resolution, the user's network, a CDN or edge POP, TLS termination, your load balancer or ingress, the service mesh sidecar, and finally the process. **Every measurement point is blind to everything upstream of it.** Choosing a vantage point is therefore choosing which class of failure your service level silently ignores. ## In-process application counters The cheapest option: increment a success or failure counter inside the handler. It is also the most dangerous, because it has a structural blind spot — **a process that is down emits no events at all.** The ratio during a total outage is not zero; it is undefined, and most systems render undefined as "no data" or carry forward the last good value, so the indicator looks healthy at the exact moment it should be screaming. It is also blind to requests rejected before reaching the handler: connection refusals when the listen backlog is full, requests dropped at the mesh sidecar, and load-balancer 502/503 responses when no healthy backend exists. All of those are unambiguous user-visible failures that never appear in the process's own numbers. Use in-process counters for internal diagnostics, not as the committed indicator. ## Load-balancer or ingress logs This is the usual default and, for most services, the right one. The load balancer is outside the failure domain of any single instance, so it keeps logging while backends die; it records 5xx responses it generated itself; it records the requests it could not place. It is close to the user while still sitting inside infrastructure you control and can be held accountable for. What it misses: DNS failures, an expired certificate at a layer in front of it, CDN or edge outages, regional network partitions between the user and your entry point, and anything that goes wrong after the bytes leave — a response that arrives but renders as an error in the client. It also cannot see a request the client never managed to send. ## Client-side telemetry Browser real-user monitoring or a mobile SDK reporting outcomes is the most faithful representation of experience: it sees the whole path, including DNS, the edge, and rendering. It is also the hardest to commit to. - **Survivorship bias.** A client that cannot reach your service frequently cannot deliver its telemetry either, so the very worst events under-report. Beacon-style delivery and retry on next launch mitigate but do not remove this. - **Faults you cannot fix.** A user on a failing mobile network or a captive portal produces failures indistinguishable from yours. You end up either accepting them into the number or building exclusion rules that are themselves contentious. - **Version lag.** Old app versions in the field keep reporting against an old definition for months. The common resolution is to keep a client-side indicator as an *informational* signal that catches classes of failure nothing else can see, while the committed target is stated at a vantage point you control. ## Synthetic probes A prober issuing a scripted journey on a schedule from several locations. Its strengths are exactly the other options' weaknesses: it exercises DNS, TLS and the edge, it works from outside your network, and it produces a constant heartbeat regardless of organic traffic — which is the decisive argument for a low-traffic or highly seasonal service where a request-ratio indicator is statistically empty overnight. Its weaknesses are resolution and representativeness. A probe every 60 seconds from five locations yields around 7,200 events a day; that cannot express a 0.5% error rate meaningfully and will not notice a failure affecting one shard, one tenant or one payload shape. Probes also tend to run the easy path with warm caches, small payloads and a test account whose data is tiny, so they systematically under-report the tail. And they are the wrong tool for anything with a side effect — you cannot probe checkout a thousand times a day without special handling. ## Making the choice The working rule: **measure as close to the user as you can still be held accountable for**, then cover the remaining blind spots with a second signal that is *not* part of the committed number. A typical composition for a public web service: - Committed availability and latency from **ingress logs**, because they are user-proximate, complete, and defensible in a review. - A **synthetic probe** journey covering DNS, certificate validity and the edge, treated as its own indicator rather than merged into the request ratio. - **Client-side** telemetry as an informational indicator, watched for divergence — when the client number is materially worse than the ingress number, something between you and the user is broken and nothing else will tell you. Two further disciplines matter. First, **do not mix vantage points inside one ratio**: summing probe events with organic requests weights them arbitrarily and makes the indicator uninterpretable. Second, keep the numerator and denominator at the *same* point — counting failures at the load balancer against a request total from the application is a classic way to manufacture an availability number nobody can reproduce. ## The tell in an interview The question separates people who have computed one of these from people who have read about them. The lived detail is the in-process counter reading 100% during an outage, and the follow-on realisation that you cannot detect "we served nothing" from a source that only speaks when it serves something.
- Your load-balancer-based availability SLI reads 99.99% but support tickets say the site is down for a whole region. What do you check?Everything upstream of the load balancer, since it is blind there: DNS resolution for that region, certificate validity, the CDN or edge POP, and network paths between users and your entry point. A synthetic probe run from inside the affected region and the client-side telemetry are the two signals that would show the gap, which is exactly why they are kept alongside the committed indicator.
- When would you make a synthetic probe the primary source for an availability SLI?When organic traffic is too sparse to compute a meaningful ratio — an internal or seasonal service that sees a few requests an hour overnight. A probe gives constant, evenly spaced events. Accept the consequences: coarse resolution, an easy path that misses tenant-specific and payload-specific failures, and the need to keep probe traffic out of the organic ratio.
- Should client-reported failures caused by the user's own network count against your SLI?Usually not against the committed target, because you cannot act on them and including them makes the number reflect carrier quality rather than your service. But do not discard them — track them separately, since a rising rate can reveal a real problem at your edge, a bad CDN configuration, or a client release that handles flaky networks badly.
saying these in an interview costs you the question
- In-process counters are the most accurate source
- Load-balancer logs see the entire user request path
- A synthetic probe every minute can measure a small error rate
- Mix probe results and real traffic into one ratio
- Client telemetry reports every failure users experience