skip to content

When replicating between Kafka clusters in different VPCs or cloud accounts, why is the network path (PrivateLink / VPC peering) a Kafka-specific challenge, and how does advertised.listeners interact with it?

level: seniorimportance: should knowfreq 40%

answer

  1. Bootstrap then direct-to-leader = two phases
  2. advertised.listeners must be routable from remote VPC
  3. Peering = L3 flat, non-overlapping CIDR + DNS
  4. PrivateLink = L4 endpoint, per-broker unique port
  5. Open the whole broker port range, not just bootstrap

basics

~20 s

Kafka clients first hit a bootstrap broker, which returns the addresses in advertised.listeners; the client then reconnects directly to each broker. Over PrivateLink/peering those advertised addresses must be resolvable and routable from the remote VPC, or the bootstrap succeeds but all follow-up connections fail.

solid answer

~50 s

Kafka uses a two-phase connection: a client connects to a bootstrap address, the broker returns metadata listing every broker by the host/port in its `advertised.listeners`, and the client then opens a **direct** connection to the specific partition leader. This breaks across private network paths if `advertised.listeners` hands back private IPs or internal hostnames that aren't routable/resolvable from the remote VPC. **VPC peering** gives flat L3 connectivity, so non-overlapping CIDRs plus DNS resolution of the advertised hostnames usually works. **AWS PrivateLink** is different: it's an L4 endpoint-service model where the consumer side gets a per-broker port via an endpoint, so each broker needs a **distinct advertised port** and the cluster must advertise the endpoint's DNS names — vanilla per-broker private IPs won't traverse PrivateLink. The practical pattern is a dedicated cross-cluster listener with advertised hostnames that resolve correctly from MM2's side, often via private hosted zones. Plus security groups/ACLs must allow the full broker port range, not just the bootstrap port.

go deeper

for a junior

Know clients reconnect to brokers using advertised.listeners, so those addresses must be reachable from the other side.

for a middle

Distinguish listeners vs advertised.listeners and that the full broker port range must be open across the link.

for a senior

Contrast peering (L3, CIDR/DNS) vs PrivateLink (L4, per-broker ports, endpoint DNS) and tie advertised hostnames to TLS SANs.

for a principal

Design the cross-account network + DNS + cert topology for many clusters, choosing peering/PrivateLink/Transit Gateway per isolation and compliance constraints.

## Kafka's two-phase wire protocol Understanding the network problem requires understanding how a Kafka client connects: 1. The client is given `bootstrap.servers` — one or a few broker addresses. 2. It connects to a bootstrap broker and asks for **metadata**: the list of all brokers, which broker leads which partition. 3. The broker answers with each broker's **advertised** address — taken from that broker's `advertised.listeners` config. 4. The client then opens a **new, direct TCP connection** to the leader of each partition it needs, using those advertised addresses. Key consequence: **the bootstrap connection working tells you nothing about whether the real data connections will work.** If advertised addresses aren't reachable from the client, you get a healthy bootstrap and then timeouts/`disconnected` on every produce/fetch. ## listeners vs advertised.listeners - `listeners` = the sockets the broker actually binds locally (e.g. `INTERNAL://0.0.0.0:9092`, `EXTERNAL://0.0.0.0:9093`). - `advertised.listeners` = what the broker *tells clients to use* per listener name. This is what gets returned in metadata. For cross-cluster replication you typically add a dedicated listener (e.g. `REPLICATION`) whose advertised value is reachable from the other cluster. ## Why VPC peering and PrivateLink differ **VPC peering** creates flat **Layer-3** routing between two VPCs. As long as the **CIDR ranges don't overlap** and DNS resolves the advertised hostnames to the right private IPs, a client in VPC-B can reach each broker's private IP directly. The advertised hostnames must resolve from VPC-B (often via a **private hosted zone** or shared Route 53 resolver). Peering is **not transitive**, so hub-and-spoke topologies need a Transit Gateway. **AWS PrivateLink** is a **Layer-4** endpoint-service model, not flat routing. The provider exposes a Network Load Balancer behind an **endpoint service**; the consumer creates an **interface endpoint** that gives them ENIs/DNS names in their own VPC. Critically, traffic is **port-mapped per target**. Because the consumer reaches brokers through endpoint DNS, the cluster must: - Advertise **endpoint DNS names**, not raw broker private IPs. - Give **each broker a unique port** on a shared advertised host (or unique hostnames), because the client must be steered to a specific broker through the single endpoint surface. This is the classic 'each broker needs its own port' PrivateLink requirement. Non-overlapping CIDRs are **not** required with PrivateLink (a reason to prefer it when address space collides), but the per-broker addressing scheme must be designed up front. ## Security groups / NACLs Whichever path you choose, firewalling must allow the **entire broker port range** used by the replication listener, not just the bootstrap port — a frequent cause of 'bootstrap works, fetch hangs'. Also allow ephemeral return traffic and the PrivateLink/peering health checks. ## Security implications (why this is in a *security* topic) PrivateLink/peering keeps replication traffic **off the public internet** entirely, shrinking the attack surface and often satisfying compliance that forbids public broker exposure. It complements (does not replace) TLS + authn/authz: the private path provides network isolation, while mTLS/SASL still authenticate identities and ACLs still authorize. Defense in depth. ## Edge cases - **Overlapping CIDRs** rule out peering; PrivateLink or NAT becomes necessary. - **DNS split-horizon**: the advertised hostname must resolve to the *peered/endpoint* address from MM2's vantage point, which may differ from how the same name resolves locally. - **Hostname verification** in TLS means the advertised hostname must also be present in the broker cert's SAN — easy to miss when introducing endpoint DNS names. - **MTU/MSS**: some peering/transit setups need MSS clamping or large Kafka messages fragment poorly.

  • The bootstrap connection works but every fetch/produce times out across PrivateLink. What's the usual cause?
    advertised.listeners returns addresses (e.g. raw broker private IPs or per-broker ports) that aren't reachable through the PrivateLink endpoint from the remote side, or the security group only opens the bootstrap port. The direct-to-leader connections fail even though bootstrap succeeded.
  • Why might you choose PrivateLink over VPC peering for cross-cluster replication?
    PrivateLink avoids needing non-overlapping CIDRs, is non-transitive-safe (unidirectional, consumer-initiated), exposes only the specific service rather than the whole VPC, and keeps traffic off the public internet — tighter isolation when address spaces collide.

saying these in an interview costs you the question

  • Assuming a successful bootstrap means data connections will work
  • Confusing listeners (bind) with advertised.listeners (what clients use)
  • Thinking PrivateLink works with raw per-broker private IPs and no per-broker port scheme
  • Forgetting that advertised hostnames must be in the broker cert SAN for TLS hostname verification
  • Treating the private path as a replacement for TLS/ACLs rather than an additional layer

context