skip to content

An availability commitment quotes a percentage over a measurement window — what exactly is measured, and what does the window hide?

level: middleimportance: must knowfreq 58%

answer

  1. a ratio, over a fixed window
  2. eligible good over eligible total
  3. unavailable means total failure
  4. latency is usually not measured
  5. averaging hides one severe outage

basics

~20 s

The figure is a ratio of eligible time or eligible requests the provider's own instrumentation judged available, over the eligible total in a fixed window — usually a calendar month, per service and per region. Averaging over that window hides short, severe outages.

solid answer

~40 s

Two definitions do all the work. First, what counts as **unavailable**: usually total failure of requests to a resource over some minimum interval, not elevated error rates and rarely latency at all — so a day of degradation your users hated can register as fully available. Second, the **measurement window**: almost always a calendar month, per service, per region, and reset at the boundary. Availability is then eligible-good over eligible-total inside that window, computed from the provider's own telemetry rather than yours. Averaging is what hides things: a single severe outage that consumes most of the month's tolerated downtime still leaves the month compliant, and an outage that straddles the month boundary is split across two windows, each of which may pass on its own.

go deeper

for a junior

Recall that the figure is a ratio over a fixed window, normally a calendar month, computed from the provider's own measurements rather than from yours. Knowing it is an average over a month is most of the answer at this level.

for a middle

Explain both definitions: what goes in the denominator, and what the contract means by unavailable — usually total failure over a minimum interval, per resource or per region, with degradation and slowness typically outside it.

for a senior

Show why a compliant month can still be a terrible month, and describe what you instrument on your own side so that you know you were affected and can evidence it later, independent of the provider's monthly figure.

for a principal

Take a position on which measurement your organisation trusts for which purpose, and where the two must be reconciled — the provider's for billing and claims, your own for engineering and for the promise you make onward to customers.

## The numerator and the denominator Every availability commitment reduces to a fraction, and the interesting content is in how each side is defined. The **denominator** is the eligible total in the window — total minutes, or total requests you actually sent. Time excluded by a carve-out is removed from this side too, not just from the bad count, which is why a heavily maintained service can post a good figure. The **numerator** is the eligible total minus what the provider counted as unavailable. That word carries a precise contractual definition, and it is usually far stricter than a user's sense of "broken": - it typically requires **total failure**, not a raised error rate — a period where some requests succeed often does not count at all; - it usually requires a **minimum duration**, so brief failures are not counted; - it is usually measured **per resource or per region**, so a failure confined to one deployment may not move a service-wide figure; - **latency** is rarely part of an availability definition; a service returning correct responses very slowly is available by the contract's lights. And the measurement is the **provider's own**. Your client-side numbers include the network path between you and the region, retries, your own timeouts and your own bugs, none of which the provider accepts. Disagreement between your dashboard and their figure is the normal case, not a sign that someone is lying. ## What the window does The measurement window is almost always a fixed calendar month. That choice has three consequences worth being able to state in an interview. 1. **It averages.** Availability is a ratio over the whole window, so the commitment implicitly tolerates a fixed amount of downtime — call it `T` minutes per window. Whether those minutes arrive as one brutal outage or as hundreds of blips is invisible in the result, even though the business impact of the two is completely different. 2. **It resets.** A bad stretch at the end of one month and another at the start of the next are scored separately, and each window may pass on its own while the fortnight spanning them was awful. 3. **It sets the granularity of the remedy.** Because the credit is computed per window, a shortfall is not a running tally; you either land below the commitment for that month or you do not. | Window property | What it gives the provider | What it costs you | |---|---|---| | Fixed calendar month | A clean billing and claim boundary | Outages spanning the boundary are split | | Ratio over the whole window | Tolerance for one long incident | One severe outage can pass while feeling catastrophic | | Per service, per region | A scope that matches how it operates | A multi-service failure is scored separately per service | ## Why compliant months can still be bad months Put the two definitions together and the failure mode is obvious. Suppose the commitment tolerates `T` minutes of downtime in a month. An outage of nearly `T` minutes in one morning is compliant. A whole day of elevated errors — some requests failing, most succeeding — may register as zero downtime, because partial failure is outside the definition. A whole day of responses so slow that clients time out may also register as zero, because latency is not usually measured. None of this is the provider cheating; it is the contract being read correctly. This is precisely why you measure your own service from your own vantage point and never treat the provider's monthly figure as a health signal. The provider's figure answers a billing question. Yours answers an engineering one. ## What to do with this - **Read the definition of unavailable before the headline figure.** It determines far more about what the commitment is worth than the number does. - **Instrument your own path.** Keep evidence with timestamps; a claim later depends on it, and so does knowing you were affected at all. - **Design against the tolerated downtime, not the percentage.** Convert the commitment into "this many minutes a month are allowed" and ask what your service does during them. - **Expect your numbers to be worse.** Yours include the network, your retries and your own failures; theirs do not. The gap is structural. The short version an interviewer wants: availability is eligible-good over eligible-total, inside a fixed monthly window, measured by the provider, under a strict definition of unavailable — and averaging across the window is what lets a painful month come out compliant.

  • Why does your own dashboard usually show worse availability than the provider's figure?
    Your view includes the network path to the region, your client timeouts, your retries and your own defects, none of which the provider accepts as its downtime. Their instrumentation also applies a strict definition of unavailable — typically total failure over a minimum interval — while your dashboard counts every failed request. The gap is structural rather than a dispute.
  • What changes if the window is a rolling period instead of a calendar month?
    A rolling window removes the boundary artefact, so an outage spanning month end is scored as one stretch rather than split into two passing windows. It also means compliance is continuously evaluated instead of settled at a billing date, which complicates a credit that is computed per invoice — one reason contractual commitments usually stay on calendar windows.

saying these in an interview costs you the question

  • Assumes the figure counts every failed request
  • Thinks slow responses count as downtime
  • Believes the measurement comes from the customer's telemetry
  • Reads the percentage without reading the definition of unavailable
  • Expects a partial-failure day to breach the commitment