skip to content

Which downtime does a provider's availability commitment typically exclude from its own measurement, and why does that matter to your contract?

level: seniorimportance: should knowfreq 40%

answer

  1. the exclusion list carries the meaning
  2. maintenance, tenant fault, preview, suspension
  3. outside provider control reaches far
  4. excluded time leaves the eligible total
  5. your customers still felt all of it

basics

~20 s

Announced maintenance, downtime caused by tenant configuration or code, preview and trial tiers, suspension for policy or non-payment, and causes outside the provider's control are normally carved out — so measured downtime is a subset of what your customers actually felt.

solid answer

~50 s

The carve-out list is where most of the commitment's real content lives. Typical exclusions: time inside an announced **maintenance window**; unavailability caused by your own configuration, code, or by exceeding a quota; anything on a preview, trial or unsupported tier; suspension under the acceptable-use or payment terms; and factors the provider says are outside its control, which usually includes the network path between your users and the region. Excluded time is normally removed from the denominator as well as the bad count, so it does not merely fail to count against the provider — it improves the ratio. The consequence when you are drafting your own customer commitment is direct: your users experience excluded downtime exactly as they experience any other, so the figure you can honestly promise has to be built on felt downtime, not on the platform's measured downtime.

code

pseudocode · 15 lines
pseudocode
totalMinutes = minutesIn(measurementWindow)
excludedMinutes = 0
for each minute in measurementWindow:
    if minute is inside an announced maintenance window:
        excludedMinutes = excludedMinutes + 1
    else if outage cause for that minute is tenant configuration:
        excludedMinutes = excludedMinutes + 1
    else if the affected tier is preview or trial:
        excludedMinutes = excludedMinutes + 1

eligibleMinutes = totalMinutes - excludedMinutes
downMinutes = count of eligible minutes the provider recorded as unavailable

measuredAvailability = (eligibleMinutes - downMinutes) / eligibleMinutes
feltDownMinutes = downMinutes + excludedMinutes    // what your customers experienced

go deeper

for a junior

Recall that exclusions exist and name the obvious ones: announced maintenance and downtime you caused yourself. Knowing the commitment does not cover everything a user experienced is the level-appropriate answer.

for a middle

Explain the full family — maintenance, tenant fault, preview tiers, suspension, causes outside the provider's control — and that excluded time is removed from the eligible total rather than counted as good time.

for a senior

Show the operating consequence: which of these you can engineer away, what evidence a claim would need, and why the gap between measured and felt downtime is the number your own service actually lives with.

for a principal

Turn it into the commitment you publish. Decide which supplier exclusions you mirror in your own contract, which you absorb deliberately, and how the residual exposure is sized before a figure is offered to customers.

## The standard carve-outs The headline figure is the part everyone quotes; the exclusion list is the part that determines what the figure means. The families are consistent across providers even where the wording differs: 1. **Announced maintenance.** Time inside a maintenance window the provider published in advance, within whatever notice period the contract states. Some services give you a choice of when the window falls; that choice is a scheduling convenience, not an exemption from the carve-out. 2. **Tenant-caused unavailability.** Downtime attributable to your configuration, your code, your credentials, or to exceeding a quota or limit you were told about. A firewall rule of yours that blocks the service is your outage, contractually. 3. **Unsupported or preview tiers.** Features in preview, trial capacity, and deployment shapes outside the committed configuration are excluded outright. 4. **Suspension under the terms.** Downtime while the account is suspended for non-payment or for an acceptable-use violation is not the provider's failure. 5. **Causes outside the provider's control.** This clause reaches further than people expect. It commonly covers the public network path between your users and the region, meaning a provider whose region is serving perfectly may be measured as fully available during an event where none of your users could reach it. ## Carve-outs move the denominator, not just the numerator A detail worth being precise about: excluded periods are normally removed from the **eligible total** as well as from the counted downtime. The measured ratio is therefore computed over a shorter window that excludes the bad parts entirely, rather than counting them as good time. The practical effect is that a heavily maintained service can post an excellent measured figure while being unavailable to users for a materially longer period than the figure implies. Nothing improper is happening; the contract is being applied as written. | Downtime your users felt | Counted against the commitment? | |---|---| | Provider incident in the region | Yes | | Announced maintenance window | Normally no — excluded from eligible time | | Your misconfigured network rule | No — tenant-caused | | A feature still in preview | No — unsupported tier | | Users' network path to the region broken | Usually no — outside provider control | | Slow but successful responses | Usually no — outside the definition of unavailable | ## Why this decides what you can promise onward This is the reason the carve-out list belongs in a design review rather than only in a procurement file. When you draft the availability commitment for your own multi-tenant reporting service, your customers measure you on the downtime they **felt**. They do not care that a period was carved out of your supplier's measurement; from their side it was an outage of your product. So the arithmetic you can honestly do is over felt downtime, and the platform's carve-outs land on your side of the line. Three practical moves follow: - **Mirror what you can.** If your suppliers exclude announced maintenance, your own commitment probably needs a maintenance-window clause too — otherwise you have absorbed a liability with no matching protection. - **Engineer around what you cannot.** Tenant-caused and preview-tier exclusions are within your control: keep production off preview features, and treat your own misconfiguration as a class of failure to be prevented rather than claimed for. - **Budget the residue.** Whatever remains — chiefly the network path to your users and the excluded maintenance you cannot pass on — is downtime you will pay for out of your own promise. Size it before you commit to a figure, not afterwards. ## How to read a commitment quickly Given a document and ten minutes, read in this order: the definition of unavailable, the exclusion list, the eligibility conditions, then the figure. Most engineers do it in reverse and form their opinion on the least informative part. An interviewer asking this question is usually checking whether you have ever actually opened one — the tell is whether you can name the tenant-caused and maintenance carve-outs without prompting, and whether you know they shrink the eligible total rather than merely failing to count.

  • Does letting you choose when the maintenance window falls change whether it is excluded?
    No. Choosing the timing moves the disruption to a less painful hour, which is worth having, but the period stays carved out of the measurement either way. It is a scheduling control, not a contractual one — and if you never pick a window, the provider chooses for you and the exclusion still applies.
  • Your users cannot reach a healthy region because of a public network event. Who owns that contractually?
    Usually nobody you can claim from. The provider measures inside its own boundary and excludes causes outside its control, so a region serving normally is measured as available. The event is real downtime for your product, which is why the path to your users belongs in your own availability plan rather than in a supplier's clause.
  • How should carve-outs change the commitment you offer your own customers?
    They set the floor of what you can absorb. Mirror the exclusions you cannot engineer away — most obviously announced maintenance — and eliminate the ones you can, by keeping production off preview tiers and treating your own misconfiguration as a defect class rather than a claimable event. Size the residue before you publish a figure.

saying these in an interview costs you the question

  • Assumes any downtime the user felt counts against the commitment
  • Thinks announced maintenance breaches an availability commitment
  • Believes excluded time still sits in the denominator
  • Runs production on a preview tier expecting a commitment
  • Expects a claim for downtime caused by own misconfiguration