skip to content

An SLO target must be evaluated over a window. Compare a rolling 28-day window with a calendar-quarter window for an availability SLO, and say what each choice changes for the team.

level: middleimportance: should knowfreq 45%

answer

  1. does the counter ever reset?
  2. four weeks holds the weekday mix
  3. week one versus week twelve
  4. internal versus contractual audiences
  5. short is jumpy, long is slow

basics

~20 s

A rolling window recomputes continuously, so an outage ages out gradually and there is never a reset to wait for. A calendar window aligns with billing and planning cycles but resets abruptly, so an early outage poisons the whole period and a late one barely registers.

solid answer

~50 s

A **rolling window** — typically 28 days, recomputed every evaluation — means the effect of an incident decays out exactly one window later and never disappears at a stroke. That keeps the picture continuous and makes it impossible to wait out a bad period, but it also means an outage keeps counting against you for four weeks after it is fixed, which can feel punitive. A **calendar window**, usually a month or a quarter, aligns with contracts, billing and planning, and gives a clean reset that management understands. Its weakness is edge effects: an outage in week one leaves the team formally in violation for the rest of the quarter with nothing to be gained, while an identical outage in the final week costs almost nothing before the counter clears. Most teams run rolling windows for internal SLOs, because they drive day-to-day decisions, and reserve calendar windows for SLA reporting, where the billing period dictates the boundary.

go deeper

for a junior

Know that an SLO is incomplete without a window, and be able to say what "99.9% over the last 28 days" means in plain words. Recognising that the window can either slide or reset is enough at this level.

for a middle

Explain the mechanics of both: continuous recomputation and gradual ageing-out versus a hard reset at a boundary. Be ready to say why 28 days is preferred over 30 for a rolling window.

for a senior

Name the edge effects with a concrete case — an outage in week one of a quarter leaving the target unreachable for eleven weeks and therefore useless as a decision input. Explain why internal SLOs and contractual reporting usually run on different windows.

for a principal

Own the incentive design: a resetting window creates a moment where risk is cheapest, and a rolling one never lets a team draw a line under an incident. Be ready to argue how you keep window changes honest so the shape of the window is never used to change a verdict.

## What the window actually is An SLO is only complete when it names a window: "99.9% of valid requests succeed" is meaningless until you say over what period. The window is the denominator of the whole arrangement, and the choice between a rolling and a calendar one changes the behaviour it encourages more than most engineers expect. ## Rolling windows A rolling window is recomputed at every evaluation over the last N days — "the trailing 28 days", continuously updated. Its properties: - **No reset to wait for.** The number today reflects the last four weeks whatever the calendar says. A team cannot sit out a bad period hoping for a fresh start. - **Gradual decay.** An incident's contribution disappears exactly N days after it happened, sliding out of the back of the window. This gives an odd effect worth naming in an interview: the measurement improves on a day when nothing was fixed, purely because an old outage aged out. - **Continuous comparability.** Any two evaluations cover an equal-length span, so "we are worse than last week" is a fair statement. ## Why 28 days and not 30 The conventional rolling length is 28 days rather than 30 because 28 is exactly four weeks. Traffic has strong weekly seasonality — weekdays and weekends have different volume and different failure profiles — and a 28-day window always contains four of each weekday, so the traffic mix is constant from one evaluation to the next. A 30-day window contains two extra days whose identity rotates, which makes the SLI drift slightly for reasons that have nothing to do with the service. ## Calendar windows A calendar window runs from the start to the end of a fixed period — a month or a quarter — and resets at the boundary. Its properties: - **Alignment with the business.** Contracts and invoices run on calendar months, so SLA reporting is nearly always monthly. Quarterly windows line up with planning cycles, which makes the SLO an input to what the team commits to next quarter. - **A clean reset.** There is a moment when the slate is genuinely clear, which is easier to communicate to non-engineers than a number that keeps drifting. - **Edge effects, in both directions.** A serious outage in the first week of a quarter can put the target out of reach for the remaining eleven weeks; once it is unattainable the target stops informing any decision, and the team spends most of the period formally in violation with nothing to do about it. The mirror problem is the end of the window: an outage in the last few days costs almost nothing because the counter is about to clear, and a team that knows this may — consciously or not — take more risk immediately after a reset. - **Incomparable periods.** February and a 92-day quarter are not the same size, so period-over-period comparisons need care. ## How teams usually resolve it The common arrangement is to run both, for different audiences. Internal SLOs use a rolling window because they drive operational decisions that happen daily and should not depend on which side of a boundary today falls. SLA compliance is reported on the calendar period the contract names, because the remedy is a credit against a specific invoice. ## Choosing the length Window length is a separate decision from rolling versus calendar, and it is a trade between responsiveness and stability: - **Short windows (a few days)** react quickly and are noisy. A single bad hour dominates them, which makes them useful for detecting a change in trend and poor for deciding policy. - **Long windows (a quarter, a year)** are stable and slow. A regression that started last week is barely visible, and by the time the number moves, the cause is cold. Four weeks is the usual compromise: long enough that a single incident does not define it, short enough that a real regression shows up within days. Very long windows have one more trap — a low-traffic service may not accumulate enough events in a short window for the ratio to mean anything, and stretching the window to fix that also stretches how long you stay blind to a regression.

  • With a rolling 28-day window, the number improves on a day when nothing was fixed. Why?
    Because an old incident just aged out of the back of the window. Rolling windows recover by the passage of time as well as by repair, which is expected behaviour but a real communication hazard — a team can appear to be recovering when nothing has changed. Annotating the timeline with when incidents entered and left the window prevents the misreading.
  • A low-traffic internal service gets a few hundred requests a day. What does that do to the window choice?
    A short window may not contain enough events for a ratio to be meaningful — at 99.9% you need thousands of events before a single failure is not a large swing. Lengthening the window stabilises the number but also lengthens how long a regression stays invisible. For very low volumes, a time-based indicator or an explicit synthetic probe schedule is usually a better fit than a request ratio.
  • Would you ever change an SLO's window mid-quarter?
    Only deliberately and with an announcement, because changing the window changes the verdict without changing the service. Shortening it after a bad incident is effectively erasing the incident, and the pattern is easy to spot. Treat window changes like contract changes: reviewed, dated, and applied going forward rather than retroactively.

saying these in an interview costs you the question

  • Stating an SLO target without naming any window
  • Believing a rolling window forgives an outage immediately once it is fixed
  • Assuming calendar and rolling windows produce the same compliance verdict
  • Choosing 30 days without noticing it breaks weekly seasonality
  • Using a quarter-long window as the day-to-day operational signal

context