skip to content

How do windowed aggregations work in ksqlDB, and what are the differences between tumbling, hopping, and session windows?

level: seniorimportance: should knowfreq 55%

answer

  1. Tumbling = fixed, non-overlapping, one bucket
  2. Hopping = fixed size + ADVANCE BY = overlap
  3. Session = inactivity gap, can merge
  4. Event-time + GRACE PERIOD = finality
  5. Key = grouping key + window bounds

basics

~10 s

Windowed aggregations group events into time buckets before aggregating. Tumbling windows are fixed-size, non-overlapping; hopping windows are fixed-size but overlap by a smaller advance; session windows are activity-based, closing after a gap of inactivity.

solid answer

~50 s

Windowing partitions an aggregation by time so each group's count/sum is computed per window. ksqlDB supports three types via the `WINDOW` clause on a GROUP BY. **Tumbling** (`WINDOW TUMBLING (SIZE 1 MINUTE)`) splits time into fixed, contiguous, non-overlapping buckets — each event lands in exactly one. **Hopping** (`WINDOW HOPPING (SIZE 1 MINUTE, ADVANCE BY 10 SECONDS)`) uses fixed-size windows that advance by a smaller step, so windows overlap and an event can belong to several — useful for moving averages. **Session** (`WINDOW SESSION (30 SECONDS)`) groups events by activity: a window stays open while events keep arriving within the inactivity gap and closes once the gap elapses, with windows merging when late events bridge them. Windowing is driven by **event-time** (record timestamps), tolerates lateness via a configurable **GRACE PERIOD** before a window's result is final, and the resulting table key becomes a composite of the grouping key plus the window bounds. Retention of old windows is bounded so state doesn't grow forever.

go deeper

for a junior

Recognize that windowing buckets aggregations by time and name the three types.

for a middle

Distinguish tumbling vs hopping vs session and write the WINDOW clause syntax.

for a senior

Explain event-time, GRACE PERIOD finality, session merging, and that the result key includes window bounds.

for a principal

Reason about state-store/changelog cost of overlapping windows, retention sizing, lateness vs correctness SLAs, and monitoring dropped-late records.

## Why windowing exists An aggregation like `COUNT(*) GROUP BY user_id` over an unbounded stream produces an ever-growing running total — there is no natural 'end'. **Windowing** bounds aggregations in **time**, so you can answer 'how many clicks per user *per minute*'. ksqlDB runs windowed aggregations on the **Kafka Streams** runtime, keying state by `(grouping key, window)`. ### Event time, not wall-clock Windows are assigned by the record's **timestamp** (event-time). By default this is the Kafka record timestamp, but you can declare a `TIMESTAMP` column. This is why out-of-order/late data needs special handling. ### The three window types **1. Tumbling** — fixed-size, contiguous, non-overlapping. Every event falls into exactly one window. ```sql SELECT user_id, COUNT(*) FROM clicks WINDOW TUMBLING (SIZE 1 MINUTE) GROUP BY user_id EMIT CHANGES; ``` `[00:00,00:01) [00:01,00:02) …` — perfect for periodic rollups. **2. Hopping** — fixed size, but advances by a smaller `ADVANCE BY`, so windows **overlap**. One event can belong to multiple windows. ```sql WINDOW HOPPING (SIZE 1 MINUTE, ADVANCE BY 10 SECONDS) ``` Produces a window every 10s, each covering 60s. Good for sliding/moving metrics. More windows = more state. **3. Session** — **activity-based**, *not* fixed size. A session window groups events for a key that occur within an **inactivity gap** of each other; it stays open as long as events keep arriving inside the gap and **closes** after the gap passes with no events. If a late event lands between two existing sessions within the gap, those sessions **merge**. ```sql WINDOW SESSION (30 SECONDS) ``` Ideal for user-activity bursts (e.g., a browsing session). ### Lateness: GRACE PERIOD Real streams have **out-of-order** records. A window can't stay open forever waiting for stragglers. ksqlDB lets you set a **grace period**: ```sql WINDOW TUMBLING (SIZE 1 MINUTE, GRACE PERIOD 10 SECONDS) ``` After `window-end + grace`, the window is **closed/final** — later records for it are **dropped**. Smaller grace = faster final results but more dropped late data; larger grace = more correctness but longer-held state and delayed finality. (Modern ksqlDB emits updates as records arrive and treats grace as the point of finality.) ### Window retention and state The result is a windowed **TABLE** whose key is the grouping key plus the window's start/end. Old windows are eventually **expired** based on a retention bound (must be ≥ window size + grace) so the **RocksDB** state store and its **changelog topic** don't grow unbounded. You can read windowed results with pull queries by specifying the key and a `WINDOWSTART`/`WINDOWEND` range. ### Edge cases - Hopping with a tiny ADVANCE produces *many* overlapping windows → high state/CPU cost. - Session windows can **merge** retroactively, which means previously-emitted session results may be superseded. - A record before `window-end + grace` updates the window; after it, the record is dropped (you can monitor dropped-late-record metrics). - Window boundaries are aligned to the epoch, not to the first event (for tumbling/hopping).

  • What is the GRACE PERIOD and what trade-off does it control?
    It's how long after a window's end ksqlDB still accepts late, out-of-order records before declaring the window final and dropping later arrivals. Larger grace = more correctness but longer-held state and delayed finality; smaller grace = faster results but more dropped data.
  • Why can session windows produce results that later change?
    Session windows are activity-based and can merge: a late event arriving within the inactivity gap between two existing sessions bridges them into one, superseding the earlier per-session results.

saying these in an interview costs you the question

  • Saying tumbling windows can overlap (only hopping overlaps)
  • Claiming windows use wall-clock time rather than event/record time
  • Ignoring that without a grace/retention bound, window state would grow unbounded
  • Describing session windows as fixed-size

context