skip to content

Throughput Metrics, Quotas and Capacity Planning

Reading bytes-in, bytes-out and message-rate metrics for capacity work, and using quotas to bound a client's share. Interviewers ask how you turn these numbers into disk and network sizing.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What do the BytesInPerSec, BytesOutPerSec, and MessagesInPerSec broker metrics measure, and where do you find them?

level: juniorimportance: must knowfreq 70%

answer

  1. kafka.server:type=BrokerTopicMetrics
  2. In=produce, Out=consumer fetch, Messages=records/s
  3. Meter: Count + 1/5/15-min rates
  4. per-broker and per-topic, not per-partition
  5. BytesOut excludes replication by default

basics

~20 s

They are JMX rate meters on the broker: BytesInPerSec is bytes produced into the broker per second, BytesOutPerSec is bytes consumers fetch out per second, and MessagesInPerSec is records produced per second. They exist both per-broker and per-topic.

solid answer

~30 s

These are Kafka's core throughput rate metrics, exposed via JMX under kafka.server:type=BrokerTopicMetrics. BytesInPerSec counts incoming producer bytes/sec, BytesOutPerSec counts bytes/sec served to consumers on fetch (not replication by default), and MessagesInPerSec counts records/sec written. Each is a Meter, so you get OneMinuteRate, FiveMinuteRate, FifteenMinuteRate, and a cumulative Count. They are available aggregated across the broker (no topic= attribute) and broken out per topic (topic=<name>). BytesInPerSec measures the compressed on-wire size as stored, so it is what drives disk and network capacity. Replication traffic is tracked separately by ReplicationBytesInPerSec/ReplicationBytesOutPerSec. Use them as the first signal for hot brokers, skewed topics, and sizing.

go deeper

for a junior

Know the three names and that In=producers, Out=consumers, Messages=record count, found via JMX.

for a middle

Know the JMX object name, Meter attributes (Count + rates), and per-topic vs aggregate views.

for a senior

Distinguish consumer egress from replication egress, know it's compressed bytes, and use it for sizing.

for a principal

Reason about leadership skew, fan-out multiplying egress, and which metric anchors disk vs network capacity models.

Kafka brokers publish operational metrics through JMX (Java Management Extensions), a standard Java mechanism for exposing runtime numbers that tools like Prometheus (via jmx_exporter), Grafana, or Datadog can scrape. **The three metrics** all live under the JMX object name `kafka.server:type=BrokerTopicMetrics,name=<MetricName>`: - **BytesInPerSec** — the number of bytes per second written into the broker by producers. Crucially this is the *compressed, on-disk* byte size (the batch as it arrives on the wire and is stored), not the uncompressed logical payload. This is the metric that most directly drives disk fill rate and inbound network. - **BytesOutPerSec** — bytes per second sent out to *consumers* via fetch requests. By default this does NOT include the bytes replicas fetch from the leader; replication has its own metrics. A single message can be counted in BytesOut many times if many consumer groups read it (fan-out). - **MessagesInPerSec** — records (messages) per second written. Note: this counts records, and with batching one produce request carries many records, so this is decoupled from request rate. **Meter semantics.** Each is a Yammer/Dropwizard `Meter`. That means each exposes four attributes: `Count` (monotonic total since broker start), and three exponentially-weighted moving averages `OneMinuteRate`, `FiveMinuteRate`, `FifteenMinuteRate`. For dashboards you almost always read the OneMinuteRate or compute your own rate by differencing `Count`. **Aggregate vs per-topic.** When the object name has no `topic=` key it is the broker-wide aggregate. With `topic=<name>` it is that topic only on that broker. There is no per-partition flavor of these; for partition-level you look at log size / lag metrics instead. **What is NOT here.** Replication traffic is `ReplicationBytesInPerSec` / `ReplicationBytesOutPerSec`. Failed/total request rates are under `RequestMetrics`. So BytesOutPerSec alone undercounts total broker egress because it omits replication. **Why it matters for capacity.** BytesInPerSec is the anchor for disk planning (disk/day = BytesInPerSec × 86400 × replicationFactor) and for inbound NIC budgeting. BytesOutPerSec plus replication tells you outbound NIC pressure, which is usually the first network bottleneck because fan-out and replication multiply egress. **Edge cases.** Compression makes BytesIn smaller than the logical data, so don't equate it with application-level throughput. A broker that hosts more leader partitions of hot topics will show higher BytesIn/Out even with balanced cluster total — that's leadership skew, fixed by reassigning preferred leaders.

  • Does BytesOutPerSec include replication traffic between brokers?
    No. Consumer fetch egress is BytesOutPerSec; replica fetch traffic is tracked separately as ReplicationBytesInPerSec / ReplicationBytesOutPerSec. Summing only BytesOutPerSec understates real outbound network.
  • Is BytesInPerSec the compressed or uncompressed size?
    Compressed (on-wire/on-disk) size as the batch is stored. So it reflects what hits the disk and network, not the uncompressed application payload.

saying these in an interview costs you the question

  • Claiming BytesOutPerSec includes inter-broker replication traffic.
  • Saying these metrics exist per-partition (they're per-broker and per-topic only).
  • Treating BytesInPerSec as the uncompressed application data size.
  • Confusing MessagesInPerSec (records) with request rate (produce requests).

context

open as a page

How do you estimate disk capacity for a Kafka cluster from throughput metrics and retention settings?

level: middleimportance: must knowfreq 60%

basics

~10 s

Disk needed = ingress bytes/sec × retention seconds × replication factor, summed across topics, plus headroom. Use BytesInPerSec for the rate and retention.ms (or retention.bytes) for how long data is kept.

open as a page

How do throughput targets and metrics drive how many partitions a topic should have?

level: middleimportance: should knowfreq 55%

basics

~20 s

Partition count must be at least target throughput divided by the per-partition throughput a single producer/consumer can sustain. Since one partition maps to one consumer in a group, partitions also set the max parallelism for consumers.

open as a page

What types of client quotas does Kafka support and how do byte-rate versus request-rate quotas differ?

level: seniorimportance: should knowfreq 50%

basics

~10 s

Kafka has two quota kinds: byte-rate quotas (producer_byte_rate, consumer_byte_rate) that cap MB/s per client, and request-rate quotas (request_percentage) that cap the share of broker request-handler/network thread time. They apply per client-id, user, or user+client-id.

open as a page

How do you estimate the network bandwidth a Kafka cluster needs, including replication, from throughput metrics?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Outbound bandwidth per broker ≈ consumer egress (BytesOutPerSec, multiplied by consumer fan-out) plus replication egress to followers (BytesInPerSec × (replicationFactor − 1)). Inbound ≈ producer ingress plus replica fetch traffic. Replication is often the dominant, easily-forgotten term.

open as a page