How does pattern subscription (subscribe with a regex) work in Kafka, and what are its operational caveats?
answer
- regex matched against cluster metadata
- metadata.max.age.ms default 5 min discovery lag
- new matching topic -> rebalance
- needs Describe ACL; scope regex tightly
- KIP-848 moves regex broker-side
basics
~10 ssubscribe(Pattern) matches topic names against a regex. The consumer auto-discovers new matching topics and rebalances to include them. It still uses group management, just with a dynamic topic set.
solid answer
~40 ssubscribe(Pattern.compile("orders-.*")) is a group-managed subscription whose topic set is resolved by regex against cluster metadata. The consumer periodically refreshes metadata (every metadata.max.age.ms, default 5 min) and when a newly-created topic matches, a rebalance pulls it in; deleted/non-matching topics drop out. Caveats: (1) the matching is metadata-driven, so newly created topics aren't picked up instantly — up to metadata.max.age.ms latency. (2) It needs Describe authorization on topics it can see; an overly broad regex plus broker-side topic visibility can leak topics or trigger frequent rebalances as topics churn. (3) Because the topic set changes underneath you, it increases rebalance frequency. It's useful for multi-tenant or sharded-topic patterns but should be scoped tightly. Pattern subscription is mutually exclusive with assign() and with topic-list subscribe().
go deeper
Know that subscribe() can take a regex and auto-discovers matching topics.
Explain metadata-refresh-driven discovery, the 5-min default lag, and rebalance churn.
Discuss ACL/Describe visibility, scoping the regex, and tuning metadata.max.age.ms tradeoffs.
Weigh dynamic topic patterns vs explicit lists in multi-tenant designs and account for KIP-848 broker-side resolution.
## Pattern (regex) subscription Besides subscribing to an explicit list of topic names, a consumer can subscribe to a **regular expression**: ```java consumer.subscribe(Pattern.compile("orders-.*"), new MyRebalanceListener()); ``` This is still **group-managed** (it joins a consumer group and rebalances). The difference is *how the set of topics is determined*: instead of a fixed list, the client matches the regex against the topic names it learns from **cluster metadata**. ### How discovery works - The consumer fetches **metadata** from the brokers. The set of topics it can see depends on broker config and ACLs (Describe permission). - Every `metadata.max.age.ms` (default **300000 ms = 5 min**) the client refreshes metadata. When the refresh reveals a **new topic matching the pattern**, the client triggers a **rebalance** so the group can start consuming it. Topics that disappear or stop matching are dropped. - Historically the regex was evaluated **client-side** over all visible topics. Newer Kafka (the next-gen **consumer rebalance protocol**, KIP-848) moves toward broker-side regex resolution, but client-side metadata refresh latency is still the mental model to know. ### Operational caveats 1. **Discovery latency:** a freshly created topic is not consumed instantly — it appears after the next metadata refresh, up to ~5 minutes unless you lower `metadata.max.age.ms` (which adds metadata traffic). 2. **Rebalance churn:** because the topic set is dynamic, topic creation/deletion drives extra rebalances. A noisy namespace = frequent pauses. 3. **Authorization & visibility:** the consumer only matches topics it is allowed to **Describe**. A broad regex (e.g. `.*`) combined with broad ACLs can pull in unintended topics; scope the regex tightly. 4. **Mutual exclusivity:** a Pattern subscription cannot coexist with a topic-list subscribe() or assign() on the same consumer. ### When to use it Multi-tenant systems where topics are created dynamically per tenant/shard (e.g. `events-tenant-*`), or where a logical stream is split across many physically-named topics. For a stable, known set of topics, prefer an explicit list to avoid the discovery latency and churn.
- Why might a newly created topic matching your pattern not be consumed immediately?Topic discovery is metadata-driven. The client only learns of the new topic on its next metadata refresh, which is bounded by metadata.max.age.ms (default 5 minutes). Lowering it speeds discovery at the cost of more metadata traffic.
saying these in an interview costs you the question
- Saying pattern subscription bypasses the consumer group — it is still fully group-managed.
- Claiming new topics are picked up instantly with zero latency.
- Forgetting it can increase rebalance frequency.
- Thinking you can combine a Pattern subscription with assign().