What do the BytesInPerSec, BytesOutPerSec, and MessagesInPerSec broker metrics measure, and where do you find them?
answer
- kafka.server:type=BrokerTopicMetrics
- In=produce, Out=consumer fetch, Messages=records/s
- Meter: Count + 1/5/15-min rates
- per-broker and per-topic, not per-partition
- BytesOut excludes replication by default
basics
~20 sThey are JMX rate meters on the broker: BytesInPerSec is bytes produced into the broker per second, BytesOutPerSec is bytes consumers fetch out per second, and MessagesInPerSec is records produced per second. They exist both per-broker and per-topic.
solid answer
~30 sThese are Kafka's core throughput rate metrics, exposed via JMX under kafka.server:type=BrokerTopicMetrics. BytesInPerSec counts incoming producer bytes/sec, BytesOutPerSec counts bytes/sec served to consumers on fetch (not replication by default), and MessagesInPerSec counts records/sec written. Each is a Meter, so you get OneMinuteRate, FiveMinuteRate, FifteenMinuteRate, and a cumulative Count. They are available aggregated across the broker (no topic= attribute) and broken out per topic (topic=<name>). BytesInPerSec measures the compressed on-wire size as stored, so it is what drives disk and network capacity. Replication traffic is tracked separately by ReplicationBytesInPerSec/ReplicationBytesOutPerSec. Use them as the first signal for hot brokers, skewed topics, and sizing.
go deeper
Know the three names and that In=producers, Out=consumers, Messages=record count, found via JMX.
Know the JMX object name, Meter attributes (Count + rates), and per-topic vs aggregate views.
Distinguish consumer egress from replication egress, know it's compressed bytes, and use it for sizing.
Reason about leadership skew, fan-out multiplying egress, and which metric anchors disk vs network capacity models.
Kafka brokers publish operational metrics through JMX (Java Management Extensions), a standard Java mechanism for exposing runtime numbers that tools like Prometheus (via jmx_exporter), Grafana, or Datadog can scrape. **The three metrics** all live under the JMX object name `kafka.server:type=BrokerTopicMetrics,name=<MetricName>`: - **BytesInPerSec** — the number of bytes per second written into the broker by producers. Crucially this is the *compressed, on-disk* byte size (the batch as it arrives on the wire and is stored), not the uncompressed logical payload. This is the metric that most directly drives disk fill rate and inbound network. - **BytesOutPerSec** — bytes per second sent out to *consumers* via fetch requests. By default this does NOT include the bytes replicas fetch from the leader; replication has its own metrics. A single message can be counted in BytesOut many times if many consumer groups read it (fan-out). - **MessagesInPerSec** — records (messages) per second written. Note: this counts records, and with batching one produce request carries many records, so this is decoupled from request rate. **Meter semantics.** Each is a Yammer/Dropwizard `Meter`. That means each exposes four attributes: `Count` (monotonic total since broker start), and three exponentially-weighted moving averages `OneMinuteRate`, `FiveMinuteRate`, `FifteenMinuteRate`. For dashboards you almost always read the OneMinuteRate or compute your own rate by differencing `Count`. **Aggregate vs per-topic.** When the object name has no `topic=` key it is the broker-wide aggregate. With `topic=<name>` it is that topic only on that broker. There is no per-partition flavor of these; for partition-level you look at log size / lag metrics instead. **What is NOT here.** Replication traffic is `ReplicationBytesInPerSec` / `ReplicationBytesOutPerSec`. Failed/total request rates are under `RequestMetrics`. So BytesOutPerSec alone undercounts total broker egress because it omits replication. **Why it matters for capacity.** BytesInPerSec is the anchor for disk planning (disk/day = BytesInPerSec × 86400 × replicationFactor) and for inbound NIC budgeting. BytesOutPerSec plus replication tells you outbound NIC pressure, which is usually the first network bottleneck because fan-out and replication multiply egress. **Edge cases.** Compression makes BytesIn smaller than the logical data, so don't equate it with application-level throughput. A broker that hosts more leader partitions of hot topics will show higher BytesIn/Out even with balanced cluster total — that's leadership skew, fixed by reassigning preferred leaders.
- Does BytesOutPerSec include replication traffic between brokers?No. Consumer fetch egress is BytesOutPerSec; replica fetch traffic is tracked separately as ReplicationBytesInPerSec / ReplicationBytesOutPerSec. Summing only BytesOutPerSec understates real outbound network.
- Is BytesInPerSec the compressed or uncompressed size?Compressed (on-wire/on-disk) size as the batch is stored. So it reflects what hits the disk and network, not the uncompressed application payload.
saying these in an interview costs you the question
- Claiming BytesOutPerSec includes inter-broker replication traffic.
- Saying these metrics exist per-partition (they're per-broker and per-topic only).
- Treating BytesInPerSec as the uncompressed application data size.
- Confusing MessagesInPerSec (records) with request rate (produce requests).