How do windowed aggregations work in ksqlDB, and what are the differences between tumbling, hopping, and session windows?
answer
- Tumbling = fixed, non-overlapping, one bucket
- Hopping = fixed size + ADVANCE BY = overlap
- Session = inactivity gap, can merge
- Event-time + GRACE PERIOD = finality
- Key = grouping key + window bounds
basics
~10 sWindowed aggregations group events into time buckets before aggregating. Tumbling windows are fixed-size, non-overlapping; hopping windows are fixed-size but overlap by a smaller advance; session windows are activity-based, closing after a gap of inactivity.
solid answer
~50 sWindowing partitions an aggregation by time so each group's count/sum is computed per window. ksqlDB supports three types via the `WINDOW` clause on a GROUP BY. **Tumbling** (`WINDOW TUMBLING (SIZE 1 MINUTE)`) splits time into fixed, contiguous, non-overlapping buckets — each event lands in exactly one. **Hopping** (`WINDOW HOPPING (SIZE 1 MINUTE, ADVANCE BY 10 SECONDS)`) uses fixed-size windows that advance by a smaller step, so windows overlap and an event can belong to several — useful for moving averages. **Session** (`WINDOW SESSION (30 SECONDS)`) groups events by activity: a window stays open while events keep arriving within the inactivity gap and closes once the gap elapses, with windows merging when late events bridge them. Windowing is driven by **event-time** (record timestamps), tolerates lateness via a configurable **GRACE PERIOD** before a window's result is final, and the resulting table key becomes a composite of the grouping key plus the window bounds. Retention of old windows is bounded so state doesn't grow forever.
go deeper
Recognize that windowing buckets aggregations by time and name the three types.
Distinguish tumbling vs hopping vs session and write the WINDOW clause syntax.
Explain event-time, GRACE PERIOD finality, session merging, and that the result key includes window bounds.
Reason about state-store/changelog cost of overlapping windows, retention sizing, lateness vs correctness SLAs, and monitoring dropped-late records.
## Why windowing exists An aggregation like `COUNT(*) GROUP BY user_id` over an unbounded stream produces an ever-growing running total — there is no natural 'end'. **Windowing** bounds aggregations in **time**, so you can answer 'how many clicks per user *per minute*'. ksqlDB runs windowed aggregations on the **Kafka Streams** runtime, keying state by `(grouping key, window)`. ### Event time, not wall-clock Windows are assigned by the record's **timestamp** (event-time). By default this is the Kafka record timestamp, but you can declare a `TIMESTAMP` column. This is why out-of-order/late data needs special handling. ### The three window types **1. Tumbling** — fixed-size, contiguous, non-overlapping. Every event falls into exactly one window. ```sql SELECT user_id, COUNT(*) FROM clicks WINDOW TUMBLING (SIZE 1 MINUTE) GROUP BY user_id EMIT CHANGES; ``` `[00:00,00:01) [00:01,00:02) …` — perfect for periodic rollups. **2. Hopping** — fixed size, but advances by a smaller `ADVANCE BY`, so windows **overlap**. One event can belong to multiple windows. ```sql WINDOW HOPPING (SIZE 1 MINUTE, ADVANCE BY 10 SECONDS) ``` Produces a window every 10s, each covering 60s. Good for sliding/moving metrics. More windows = more state. **3. Session** — **activity-based**, *not* fixed size. A session window groups events for a key that occur within an **inactivity gap** of each other; it stays open as long as events keep arriving inside the gap and **closes** after the gap passes with no events. If a late event lands between two existing sessions within the gap, those sessions **merge**. ```sql WINDOW SESSION (30 SECONDS) ``` Ideal for user-activity bursts (e.g., a browsing session). ### Lateness: GRACE PERIOD Real streams have **out-of-order** records. A window can't stay open forever waiting for stragglers. ksqlDB lets you set a **grace period**: ```sql WINDOW TUMBLING (SIZE 1 MINUTE, GRACE PERIOD 10 SECONDS) ``` After `window-end + grace`, the window is **closed/final** — later records for it are **dropped**. Smaller grace = faster final results but more dropped late data; larger grace = more correctness but longer-held state and delayed finality. (Modern ksqlDB emits updates as records arrive and treats grace as the point of finality.) ### Window retention and state The result is a windowed **TABLE** whose key is the grouping key plus the window's start/end. Old windows are eventually **expired** based on a retention bound (must be ≥ window size + grace) so the **RocksDB** state store and its **changelog topic** don't grow unbounded. You can read windowed results with pull queries by specifying the key and a `WINDOWSTART`/`WINDOWEND` range. ### Edge cases - Hopping with a tiny ADVANCE produces *many* overlapping windows → high state/CPU cost. - Session windows can **merge** retroactively, which means previously-emitted session results may be superseded. - A record before `window-end + grace` updates the window; after it, the record is dropped (you can monitor dropped-late-record metrics). - Window boundaries are aligned to the epoch, not to the first event (for tumbling/hopping).
- What is the GRACE PERIOD and what trade-off does it control?It's how long after a window's end ksqlDB still accepts late, out-of-order records before declaring the window final and dropping later arrivals. Larger grace = more correctness but longer-held state and delayed finality; smaller grace = faster results but more dropped data.
- Why can session windows produce results that later change?Session windows are activity-based and can merge: a late event arriving within the inactivity gap between two existing sessions bridges them into one, superseding the earlier per-session results.
saying these in an interview costs you the question
- Saying tumbling windows can overlap (only hopping overlaps)
- Claiming windows use wall-clock time rather than event/record time
- Ignoring that without a grace/retention bound, window state would grow unbounded
- Describing session windows as fixed-size