skip to content

How do cross-region/cross-AZ egress charges shape Kafka multi-region architecture, and what techniques reduce that cost?

level: seniorimportance: should knowfreq 50%

answer

  1. egress billed per region/AZ boundary
  2. RF + cross-zone reads multiply bytes
  3. KIP-392 follower fetch = read locally
  4. compression shrinks every hop
  5. replicate only topics that need it

basics

~20 s

Cloud providers charge for data leaving a region or AZ. Kafka replication and cross-region consumers multiply traffic, so egress can dominate the bill. Reduce it with follower fetching (read locally), compression, fewer cross-region replicas, and rack-aware placement to avoid needless cross-AZ hops.

solid answer

~40 s

In the cloud you pay for inter-region and (often) inter-AZ data transfer. Kafka amplifies this: a record written once is replicated to N replicas and then read by every consumer. If replicas or consumers sit in other regions/AZs, each copy and each fetch is billable egress. With a replication factor of 3 across AZs plus several cross-region consumers, raw egress can dwarf compute cost. Mitigations: (1) **Follower fetching** (KIP-392, `client.rack` + `replica.selector.class=...RackAwareReplicaSelector`) lets consumers read from a same-AZ/region follower, avoiding cross-zone reads. (2) **Producer/broker compression** (`compression.type`) shrinks bytes on the wire. (3) Keep cross-region replication to a single async stream (MirrorMaker 2 / Cluster Linking) rather than stretching synchronous replicas everywhere. (4) **Rack awareness** to control replica placement. (5) Replicate only the topics that need it, and consider tiered storage to cut re-fetch volume.

go deeper

for a junior

Know that moving data between regions/AZs costs money and Kafka moves a lot of data.

for a middle

Name follower fetching and compression as cost levers and know RF multiplies traffic.

for a senior

Model where each byte is billed (replication, reads, cross-region copy) and configure KIP-392 + rack awareness.

for a principal

Build a cost model per topic and design placement/replication to meet HA within an egress budget.

## Why egress is the hidden Kafka bill Cloud billing charges **data transfer out** of an AZ and especially out of a region (inter-region egress is far pricier; internet egress pricier still). Kafka's design multiplies bytes: - A record is **replicated** to `replication.factor` brokers. With RF=3 across three AZs, the leader sends each record to two other AZs — billable inter-AZ transfer on every produce. - Every **consumer fetch** reads from the partition leader by default; a consumer in a different AZ/region pays cross-zone/cross-region transfer for all data it reads. - **Cross-region replication** (MM2/Cluster Linking) ships entire topics across the priciest boundary. So the same record can be billed multiple times: on replication, on every cross-zone consumer read, and on cross-region copy. ## Techniques to cut it 1. **Follower fetching (KIP-392).** Historically consumers had to read from the leader. KIP-392 lets a consumer fetch from the **nearest in-sync follower**. Configure brokers with `replica.selector.class` set to the rack-aware selector and have consumers set `client.rack`. A consumer then reads from a same-AZ replica, eliminating cross-zone read egress. This is the single biggest lever for consumer-side cost. 2. **Compression.** `compression.type=zstd|lz4|gzip|snappy` at the producer (and matching broker behavior) shrinks bytes traveling across every billable boundary — replication, fetch, and cross-region copy. 3. **Limit synchronous cross-region replicas.** A fully stretched cluster sends synchronous replication traffic across regions repeatedly. Prefer one async cross-region stream and keep RF local. 4. **Rack-aware replica placement.** `broker.rack` + rack-aware assignment spreads replicas for HA but can be tuned to avoid unnecessary cross-zone hops. 5. **Replicate selectively.** Only mirror the topics a remote region actually needs; don't blanket-copy everything. 6. **Tiered storage.** Offloading old segments to object storage reduces broker disk and can cut some re-fetch transfer, though object-store egress has its own pricing. ## Edge cases / gotchas - Follower fetching only helps **reads**; it does not reduce replication egress, which is driven by RF and placement. - A follower you fetch from must be in-sync; if it falls out of ISR the consumer falls back to the leader (cost spikes). - Some managed offerings bill differently (e.g., per-partition or networking-inclusive pricing) — model the actual provider's rate card. - Compression trades CPU for bytes; very high ratios can raise broker CPU and produce latency.

  • Does follower fetching reduce replication egress as well as consumer-read egress?
    No. KIP-392 only affects where consumers read from, eliminating cross-zone/region read traffic. Replication egress is determined by replication.factor and replica placement (rack awareness), not by where consumers fetch — so you must address it separately by tuning RF, placement, and how many regions hold synchronous replicas.
  • What configs enable follower fetching?
    On the broker, set replica.selector.class to org.apache.kafka.common.replica.RackAwareReplicaSelector and assign each broker a broker.rack. On the consumer, set client.rack to its zone/region. The broker then routes fetches to the nearest in-sync follower in the consumer's rack.

saying these in an interview costs you the question

  • Assuming egress is negligible compared to compute (it often dominates Kafka cost).
  • Claiming follower fetching reduces replication traffic (it only affects consumer reads).
  • Ignoring inter-AZ charges and only thinking about inter-region.
  • Replicating every topic to every region by default.

context