skip to content

How does the circuit breaker's sliding window work — count-based vs time-based — and how are failure and slow-call rates evaluated?

level: seniorimportance: should knowfreq 40%

answer

  1. COUNT_BASED = last N calls; TIME_BASED = last N seconds
  2. two rates: failure + slow-call
  3. slow call counts even if it succeeds
  4. minimum-number-of-calls gates evaluation
  5. size = count vs seconds depending on type

basics

~20 s

The sliding window stores recent call outcomes. COUNT_BASED keeps the last N calls; TIME_BASED keeps calls from the last N seconds. The breaker computes failure-rate and slow-call-rate over the window and trips to OPEN when either crosses its threshold, but only after minimum-number-of-calls.

solid answer

~40 s

Resilience4j records each call's outcome (success / failure / slow) in a ring buffer called the sliding window. **COUNT_BASED** (`sliding-window-type: COUNT_BASED`, `sliding-window-size: N`) aggregates the last N calls. **TIME_BASED** aggregates all calls in the last N seconds, using per-second partial aggregations so cost is O(1) regardless of throughput. From the window it derives **failure-rate** (percentage of recorded failures) and **slow-call-rate** (percentage exceeding `slow-call-duration-threshold`). If either meets or exceeds its configured threshold — `failure-rate-threshold` or `slow-call-rate-threshold` — the breaker transitions CLOSED->OPEN. Crucially, the rates are only *evaluated* once `minimum-number-of-calls` outcomes exist; below that, the breaker stays CLOSED. Slow calls are counted as slow even when they succeed, so a healthy-but-degraded dependency can still trip the breaker on latency alone.

code

yaml · 20 lines
yaml
resilience4j:
  circuitbreaker:
    instances:
      # Count-based: decisions over the last 20 calls
      searchApi:
        sliding-window-type: COUNT_BASED
        sliding-window-size: 20
        minimum-number-of-calls: 10
        failure-rate-threshold: 50
        slow-call-rate-threshold: 80
        slow-call-duration-threshold: 1s
        wait-duration-in-open-state: 10s
        permitted-number-of-calls-in-half-open-state: 5
        automatic-transition-from-open-to-half-open-enabled: true
      # Time-based: decisions over the last 60 seconds of traffic
      reportApi:
        sliding-window-type: TIME_BASED
        sliding-window-size: 60     # seconds, not calls
        minimum-number-of-calls: 20
        failure-rate-threshold: 40

go deeper

for a junior

Know the window holds recent outcomes and there are count- and time-based variants.

for a middle

Explain failure-rate vs slow-call-rate and the minimum-number-of-calls gate.

for a senior

Discuss TIME_BASED per-second aggregation, latency-only tripping, and HALF_OPEN's separate trial window.

for a principal

Tune window type/size against traffic profile and SLOs, reason about statistical stability of small windows, and correlate breaker events with downstream capacity.

**Purpose of the window.** The circuit breaker needs a rolling sample of recent behavior to decide health. That sample is the **sliding window**, and its type determines *what 'recent' means*. **COUNT_BASED window.** - Config: `sliding-window-type: COUNT_BASED`, `sliding-window-size: 10`. - Implemented as a fixed-size ring buffer of the **last N call outcomes**. When call N+1 arrives, the oldest is evicted. - Aggregated metrics (total, failed, slow) are maintained incrementally so reads are O(1). - Good when throughput is steady; the window's time-span depends on traffic (10 calls could span 1s or 10 minutes). **TIME_BASED window.** - Config: `sliding-window-type: TIME_BASED`, `sliding-window-size: 60` (seconds). - Keeps all outcomes from the **last N seconds**, using `N` per-second **partial aggregations** (buckets). Each second gets a bucket; buckets older than N seconds are dropped. This keeps memory bounded and updates O(1) even under high load. - Good when you care about a fixed time horizon regardless of traffic volume. **The two rates.** 1. **Failure rate** = (failed calls / recorded calls) * 100. A 'failure' is any recorded exception except those in `ignore-exceptions`; if `record-exceptions` is set, only those count. 2. **Slow-call rate** = (slow calls / recorded calls) * 100. A call is 'slow' if its duration >= `slow-call-duration-threshold`, **regardless of success**. This lets the breaker react to degradation before hard failures appear. **Evaluation gating: `minimum-number-of-calls`.** The breaker computes and acts on the rates *only after* at least this many calls are recorded in the current window. Example: window size 10, minimum 10, threshold 50% — the breaker won't open until 10 calls exist and >=5 are failures/slow. Set too high, the breaker is sluggish; too low, it can trip on tiny samples. **Transition mechanics.** - CLOSED -> OPEN: when failure-rate >= `failure-rate-threshold` OR slow-call-rate >= `slow-call-rate-threshold` (evaluated once minimum met). - OPEN -> HALF_OPEN: after `wait-duration-in-open-state` (optionally with exponential backoff via `enable-exponential-backoff` in newer versions), or automatically if `automatic-transition-from-open-to-half-open-enabled: true` (otherwise transition happens on the next call after the wait elapses). - HALF_OPEN: allows `permitted-number-of-calls-in-half-open-state` trial calls; it uses a fresh window sized to that count and re-evaluates the same rates to decide CLOSED vs OPEN. **Gotchas.** - Confusing `sliding-window-size` units: it's a **count** for COUNT_BASED but **seconds** for TIME_BASED. - A breaker can open with **zero exceptions** purely from slow calls. - `minimum-number-of-calls` should generally be <= `sliding-window-size`; a larger minimum means the rate is never evaluated until the window is full. - The window is shared across all threads for that named instance — it reflects aggregate traffic, not per-caller. - In HALF_OPEN, calls beyond the permitted number are rejected with `CallNotPermittedException`, same as OPEN. **When to choose which.** Use COUNT_BASED for consistent, high-enough traffic where 'last N calls' is meaningful. Use TIME_BASED for low or bursty traffic where you want decisions over a wall-clock horizon rather than a raw call count.

  • Can a circuit breaker open even if no call ever threw an exception?
    Yes. If the slow-call-rate over the window reaches slow-call-rate-threshold — where a call is 'slow' when its duration exceeds slow-call-duration-threshold regardless of success — the breaker opens on latency alone, with zero exceptions recorded.
  • For a TIME_BASED window, what does sliding-window-size mean?
    It's the number of seconds of history retained, not a call count. Resilience4j keeps per-second partial aggregations for that many seconds and drops older buckets, so the rate reflects all calls in the last N seconds.

saying these in an interview costs you the question

  • Thinking sliding-window-size is always a call count (it's seconds for TIME_BASED)
  • Believing only exceptions can trip the breaker (ignores slow-call rate)
  • Assuming the rate is evaluated from the first call, ignoring minimum-number-of-calls
  • Thinking each caller/thread has its own window rather than a shared per-instance window

context