skip to content

Explain the difference between a push query (EMIT CHANGES) and a pull query in ksqlDB.

level: middleimportance: must knowfreq 72%

answer

  1. Push = EMIT CHANGES, never ends, subscribe
  2. Pull = point lookup, returns now, like GET
  3. Pull needs a materialized table
  4. Push works on streams + tables
  5. Pull routes to key's partition owner

basics

~20 s

A push query (EMIT CHANGES) subscribes to a continuous stream of updates and never finishes until cancelled — it pushes new results as they arrive. A pull query is a point-in-time lookup against a materialized table's current state and returns immediately.

solid answer

~50 s

ksqlDB serves two query modes. A **push query** uses `EMIT CHANGES` and stays open indefinitely: it emits every new/updated row as events flow through the underlying topology — ideal for live dashboards, alerting, or feeding an app a subscription of changes. It runs over both streams and tables. A **pull query** is a synchronous, request/response lookup that returns the **current value** for a key (or a key range) from a **materialized table** and then completes — like a key-value GET, used for low-latency reads from your app. Pull queries require the table to be **materialized** (created via an aggregation/CTAS so a state store exists) and historically were limited to point lookups by the table's key, with range/full-scan support added over later versions. Push queries are typically issued over the streaming HTTP endpoint; pull queries return a finite result set and integrate well with REST request/response. Both hit the ksqlDB `/query` (or `/query-stream`) endpoint.

go deeper

for a junior

Know push = EMIT CHANGES = continuous; pull = one-time lookup that returns now.

for a middle

Explain that pull needs a materialized table, push works on streams/tables, and typical use cases for each.

for a senior

Cover state-store routing across nodes, eventual consistency/lag, /query-stream, and LIMIT semantics on push.

for a principal

Discuss serving-layer design: when to expose pull queries as an app read path vs. an external KV store, and consistency trade-offs.

## Two ways to read data in ksqlDB ### Push query — `EMIT CHANGES` A **push query** turns a SQL statement into a **subscription**. You ask 'tell me every result, now and forever, as data changes,' and ksqlDB keeps the connection open, **pushing** each new or updated row to you as records flow through the underlying Kafka Streams topology. ```sql SELECT * FROM clicks WHERE url = '/checkout' EMIT CHANGES; ``` This never terminates on its own — it runs until you cancel it (or hit a `LIMIT`). Push queries can run over **streams** (every matching event) or **tables** (every change to the materialized state, i.e. the changelog). They are the basis for live dashboards, real-time alerting, and change-data subscriptions. ### Pull query — point-in-time lookup A **pull query** is a **request/response** read of the *current* state. It looks up a key (or key range) in a **materialized view** and returns immediately, like a database SELECT or a key-value GET: ```sql SELECT order_count FROM orders_per_user WHERE user_id = 'u42'; ``` No `EMIT CHANGES`. It returns the current value and the query completes. This is what your application calls when it needs a fast read of the latest aggregate. ### Why pull queries need a materialized table For ksqlDB to answer 'what is the current value for key X' instantly, it must already hold that state. A **materialized view** is created when you run an aggregation (CTAS with GROUP BY), which builds a **RocksDB-backed state store** continuously updated by the persistent query. A pull query reads directly from that state store (routing to whichever ksqlDB node owns the key's partition). You **cannot** pull from a raw stream — there is no current-state store to read. Over ksqlDB versions, pull-query capability expanded from single-key lookups to **key ranges**, table scans, and pull queries over streams in newer releases. ### Mechanics and routing - In a multi-node ksqlDB cluster, state is **partitioned**. A pull query for key X is routed (via interactive-query metadata) to the node hosting that key's partition; if local, it reads the local store, otherwise it **forwards** the request to the owning node. - Push queries fan out results from the running topology; with `/query-stream` (HTTP/2) ksqlDB streams them efficiently. ### Edge cases / gotchas - A pull query against state that is still **rebuilding** (after a restart/restore) may fail or wait until the store is caught up. - Push queries consume resources for as long as they live; many open push queries add load. - Pull queries see the **latest committed** materialized value, which reflects processing/commit lag — they are eventually consistent with the source topic, not the instantaneous head. - A `LIMIT n` on a push query makes it terminate after n rows — a bounded push, not a pull.

  • Why can't you run a pull query against a raw STREAM (in older versions)?
    A pull query reads current state from a materialized state store. A raw stream has no such store — there is no 'current value per key' to look up — so you need a materialized table (built via an aggregation/CTAS).
  • How does a pull query find the right answer in a multi-node ksqlDB cluster?
    State is partitioned across nodes. ksqlDB uses interactive-query metadata to route the lookup to the node that owns the key's partition, forwarding the request if the key isn't local.

saying these in an interview costs you the question

  • Saying pull queries stream continuously (that's a push query)
  • Claiming push queries return a finite result and exit on their own (they run until cancelled or LIMIT)
  • Asserting you can pull from any stream without materialization

context