skip to content

What does selectKey do, how does it relate to groupBy and the deprecated through(), and what is the modern replacement?

level: seniorimportance: should knowfreq 35%

answer

  1. selectKey = new key, same value, flags repartition
  2. groupBy = selectKey + groupByKey
  3. through() deprecated
  4. repartition() = Streams-owned internal topic
  5. to()+stream() = real user topic

basics

~20 s

selectKey sets a new record key (value unchanged) and flags repartition. groupBy is essentially selectKey + groupByKey. The old through() (write-to-topic-then-reconsume) is deprecated; use repartition() instead for the implicit re-keying topic, or to()+stream() for an explicit user topic.

solid answer

~40 s

selectKey((k,v) -> newKey) produces a stream with a new key and the same value; like map/flatMap it sets the repartition-required flag because the key changed. It's preferred over map when you ONLY want to rekey, since the intent is explicit. groupBy(keySelector) is sugar for selectKey(...).groupByKey() — it rekeys and prepares for aggregation, triggering a repartition topic. The deprecated through("topic") used to write to a named topic and immediately reconsume from it (a manual repartition round-trip); it's replaced by repartition() for Streams-managed internal repartition topics (auto-named or via Repartitioned.as(...)), or by explicit to("topic") + builder.stream("topic") when you want a real, user-visible topic. The reason through() was deprecated is that repartition() lets Streams own the internal topic's lifecycle, partitioning, and naming, while through forced you to pre-create and manage the topic yourself.

go deeper

for a junior

Know selectKey sets a new key.

for a middle

Relate groupBy = selectKey + groupByKey and know through() is replaced by repartition().

for a senior

Explain when to use repartition() vs to()+stream() and why through() was deprecated.

for a principal

Design rekey topology with stable internal-topic naming, partition budgeting, and contract vs internal topic boundaries.

## selectKey `KStream.selectKey((key, value) -> newKey)` returns a new `KStream` whose records have a **new key** and the **unchanged value**. It is stateless. Because it changes the key, it sets the **repartition-required flag** — identical consequence to map/flatMap, but selectKey signals intent clearly ('I'm only changing the key'). Use it when, for example, your records are keyed by event-id but you need to aggregate by customer-id: `selectKey((k,v) -> v.customerId)`. ## selectKey vs groupBy `groupBy(keySelector)` is conceptually `selectKey(keySelector).groupByKey()`. groupByKey alone keeps the existing key (and assumes no upstream rekey requiring repartition); groupBy explicitly re-keys and so always implies a repartition before the aggregation. If your data is already keyed correctly, use `groupByKey()` and avoid the rekey/repartition; only use `groupBy` when you must aggregate on a different key. ## through() — deprecated Historically, `KStream.through("my-topic")` wrote every record to `my-topic` and then immediately re-consumed from it, returning a stream reading from that topic. People used it as a manual repartition step (force records onto a topic whose partitioning matches a new key). Drawbacks: you had to **pre-create and manage** the topic, its partition count, retention, and cleanup yourself, and the round-trip semantics were implicit. ## Modern replacements - **`repartition()`** (Kafka 2.6+): `stream.repartition()` or `stream.repartition(Repartitioned.as("name").withNumberOfPartitions(n).withStreamPartitioner(...).withKeySerde(...).withValueSerde(...))`. Streams creates and **owns** an internal repartition topic — manages its lifecycle, partition count, retention, and naming. This is the direct replacement for through() when you want an internal, Streams-managed repartition. - **`to("topic")` + `builder.stream("topic")`**: when you want a **real, user-visible** topic (e.g. another app consumes it), write it explicitly with to() and re-read it with a new stream. This is the replacement when the topic is part of your contract, not just an internal repartition. ## Putting it together A typical rekey-then-aggregate pipeline: ``` builder.stream("events") .selectKey((k, v) -> v.getCustomerId()) // sets repartition flag .repartition(Repartitioned.as("by-customer")) // optional explicit control .groupByKey() .count(); ``` Without the explicit repartition(), groupByKey() after the selectKey still triggers an auto-named repartition topic. The explicit call just gives you a stable name and control over partitions/serdes.

  • When should you use to()+stream() instead of repartition() to re-key data?
    When the intermediate topic is a real, contractually-visible topic that other applications consume, rather than an internal repartition. repartition() creates a Streams-managed internal topic with an opaque lifecycle; to()+stream() gives you a normal user topic you control and others can read.
  • If your input is already keyed by the aggregation key, should you use groupBy or groupByKey?
    groupByKey — it keeps the existing key and avoids the rekey-induced repartition that groupBy would force. Use groupBy only when you must aggregate on a different key than the current one.

saying these in an interview costs you the question

  • Saying selectKey changes the value
  • Recommending through() in new code (it's deprecated)
  • Claiming groupBy and groupByKey are interchangeable with the same cost
  • Thinking selectKey alone creates a repartition topic without a downstream stateful op

context