You run a 3-AZ Kafka cluster on AWS and your bill is dominated by cross-AZ data transfer from consumers. As a principal engineer, how do you design a fetch-from-follower rollout to actually capture the savings, and what can undermine them?
answer
- FFF fixes only the consumer-read cross-AZ slice
- placement must cover every consumer AZ (RF vs #AZs)
- rack typo = silent cross-AZ, no error
- derive client.rack from k8s zone label
- measure with flow logs/cost, not config presence
basics
~20 sTag brokers with broker.rack=AZ, enable RackAwareReplicaSelector, ensure every partition has a replica in each consumer AZ (rack-aware placement), and set each consumer's client.rack to its AZ. Savings collapse if there's no in-rack replica, racks are misconfigured, or consumers fall back to the leader.
solid answer
~50 sCapture savings by aligning placement and selection. (1) Set broker.rack to the real AZ on every broker and use rack-aware replica assignment so each partition has an in-sync replica in every consumer AZ — otherwise the selector has nothing local to pick and reads stay cross-AZ. (2) Enable replica.selector.class=RackAwareReplicaSelector cluster-wide (dynamically, validated). (3) Set client.rack on every consumer to match its AZ exactly. (4) Verify with metrics — measure cross-AZ bytes before/after, not just config presence. Undermining factors: a misconfigured/typo'd rack silently sends reads cross-AZ; partitions whose replicas don't cover an AZ; ISR shrink/leader changes causing temporary leader fallback; replication-factor/AZ mismatch (RF=2 over 3 AZs can't cover all zones); and the produce path and inter-broker replication are still cross-AZ, so FFF only addresses the consumer-read slice of the bill. Also weigh the freshness cost for latency-sensitive consumers.
go deeper
Understand the goal: keep reads in the same AZ to save money; the details are for more senior folks.
Know you need both the configs and a replica present in the consumer's AZ.
Explain placement-vs-selection alignment, RF/AZ coverage, and fallback-on-churn erosion.
Architect the full rollout (rack hygiene, placement, dynamic enablement, k8s zone wiring, measurement) and articulate what FFF does and does NOT cover in the egress bill.
## Where the money goes In a multi-AZ AWS Kafka cluster there are three cross-AZ traffic sources: (a) **producer → leader** (often cross-AZ), (b) **leader → follower replication** (cross-AZ by design, for durability), and (c) **leader → consumer reads** (cross-AZ when the consumer isn't co-located with the leader). KIP-392 fetch-from-follower attacks **only (c)** — the consumer-read slice. For fan-out-heavy workloads (many consumer groups), (c) is frequently the largest line item, which is why FFF can be a big win — but it does **not** remove (a) or (b). ## Designing the rollout ### Step 1 — Honest rack labeling Set `broker.rack` on every broker to its true AZ (`us-east-1a/b/c`). This label is the locality key. A typo or stale value silently routes reads cross-AZ with no error — only the bill reveals it. Treat rack as config-managed and validated across the fleet. ### Step 2 — Placement must cover every consumer AZ The selector can only return a replica that **exists and is in-sync in the consumer's rack**. So replica **placement** must put a replica of each partition in each AZ that has consumers. Practically: use **rack-aware assignment** and pick a replication factor that covers your AZs (e.g., RF=3 across 3 AZs). RF=2 across 3 AZs leaves one AZ with no local replica for some partitions → those reads stay cross-AZ. Placement and selection are two halves of one design; either alone is insufficient. ### Step 3 — Enable the selector cluster-wide `replica.selector.class=RackAwareReplicaSelector`, applied via dynamic broker config (or rolling restart), validated on all brokers. Missing it on even some brokers leaves those leaders serving cross-AZ. ### Step 4 — Consumer client.rack Every consumer sets `client.rack` to its own AZ. In Kubernetes, derive it from the node's topology label (`topology.kubernetes.io/zone`) via the downward API so pods self-configure correctly regardless of where they're scheduled. ### Step 5 — Measure, don't assume Validate with **cross-AZ byte metrics / VPC flow logs / cost reports**, not just config. Confirm `preferredReadReplica` is being honored (consumer metrics) and watch the cross-AZ transfer line drop. ## What undermines the savings - **No in-rack replica** for some partitions → those reads remain cross-AZ. - **Rack misconfiguration** → silent cross-AZ reads, no error. - **Leader/ISR churn** → temporary fallback to the leader (cross-AZ) until consumers re-home; frequent rebalancing erodes savings. - **Hot partitions / skew** → if consumers cluster in one AZ but leaders/replicas don't, coverage gaps. - **Ignoring (a) and (b)** → produce and replication remain cross-AZ; FFF is not a total-egress solution. - **Latency-sensitive consumers** forced onto followers pay in freshness; sometimes correct answer is to keep them on the leader. ## Strategic framing FFF is one tool in cost-optimized multi-AZ design. Complementary moves: co-locating producers near leaders, accepting replication cost as the price of durability, and (for cross-*region* DR) using MirrorMaker 2 rather than stretching one cluster. FFF keeps leaders single-region while letting reads be AZ-local — exactly the 'leaders stay put, reads go local' pattern.
- If RF=2 in a 3-AZ cluster, why might FFF still leave some reads cross-AZ?Each partition has replicas in only 2 of the 3 AZs. Consumers in the third AZ have no local replica to select, so their reads fall back to a cross-AZ replica or the leader.
- Does FFF reduce the leader→follower replication cross-AZ cost?No. Replication traffic between leader and followers is unchanged — it's the price of durability. FFF only redirects consumer reads.
- How would you set client.rack reliably for consumers running on Kubernetes?Read the node's topology.kubernetes.io/zone label via the downward API and inject it as client.rack, so each pod advertises the AZ it's actually scheduled in.
saying these in an interview costs you the question
- Claiming FFF eliminates all cross-AZ cost — it leaves produce and replication traffic untouched.
- Enabling the selector but never fixing replica placement to cover consumer AZs.
- Trusting config presence instead of measuring actual cross-AZ bytes.
- Forcing latency-critical consumers onto followers without considering freshness lag.