Which metrics and signals do you monitor to verify tiered storage is healthy in production, and what does rising remote-copy lag indicate?
answer
- RemoteCopyLagBytes/Segments = primary leading signal
- upload<ingest ⇒ lag grows ⇒ local disk fills
- RemoteCopy/Fetch/DeleteErrorsPerSec for backend faults
- copier thread-pool saturation + local disk headroom
- watch __remote_log_metadata + RemoteDelete lag
basics
~10 sWatch remote copy lag (bytes/segments waiting to upload), remote copy throughput, and upload/delete error rates, plus local disk usage. Rising copy lag means uploads can't keep up, so segments stay local and disk fills.
solid answer
~40 sThe key operational signal is the RemoteLogManager copy lag: RemoteCopyLagBytes / RemoteCopyLagSegments (per topic-partition) show data rolled but not yet uploaded. Pair it with RemoteCopyBytesPerSec / RemoteDeleteBytesPerSec throughput and RemoteCopyErrorsPerSec / RemoteFetchErrorsPerSec / RemoteDeleteErrorsPerSec. Also track RemoteLogReader request rate/latency for the read path, the remote.log.manager copier/expiration thread-pool saturation, and—most importantly operationally—local disk utilization. When copy lag rises, uploads aren't keeping up with ingest (slow/erroring backend, under-sized thread pool, throttling, or credential/connectivity issues); because a segment can't be deleted locally until it's uploaded, local disk fills toward the disk-full failure. You respond by checking error metrics, increasing remote.log.manager.thread.pool.size / copier concurrency, raising backend throughput limits, or throttling ingest. Alert on sustained copy lag and on local disk headroom before it becomes critical.
go deeper
Knows to watch local disk and whether segments are uploading.
Names RemoteCopyLag and error metrics and links lag to local disk filling.
Diagnoses lag causes (throughput vs errors), tunes copier thread pools and backend limits, and watches the read path.
Defines SLOs/alerts across the fleet for copy/delete lag, metadata-topic health, and disk headroom, and runbooks the responses.
**Why monitoring matters.** Tiered storage adds an asynchronous pipeline: rolled segments are uploaded by the RemoteLogManager. If that pipeline stalls, segments accumulate locally (they can't be deleted before upload), and the broker eventually runs out of disk. So the health of tiered storage is mostly about the *copy pipeline* and *local disk headroom*. **Core metrics (Kafka JMX, under kafka.server / RemoteLogManager):** - **RemoteCopyLagBytes / RemoteCopyLagSegments** (per topic, often per-partition): how much rolled data is waiting to be uploaded. The primary leading indicator. Steady ≈ small; rising ⇒ uploads falling behind. - **RemoteCopyBytesPerSec**: upload throughput. Compare to ingest rate — if upload < ingest for tiered data, lag grows. - **RemoteDeleteLagBytes / RemoteDeleteBytesPerSec**: remote retention deletion progress; if deletes stall, remote cost grows unbounded. - **RemoteCopyErrorsPerSec / RemoteFetchErrorsPerSec / RemoteDeleteErrorsPerSec**: backend/connectivity/permission failures. Non-zero copy errors usually explain rising lag. - **Read path:** RemoteLogReader / RemoteFetchBytesPerSec and remote fetch latency, plus the remote-read thread pool (remote.log.reader.threads) queue size — high values mean cold reads are saturating. - **Thread-pool saturation:** the copier and expiration thread pools (remote.log.manager.thread.pool.size and related copier/expiration pools) — if fully busy, increase size. - **Local disk utilization per broker** — the ultimate failure mode; alert with enough headroom to react. - **__remote_log_metadata** topic health (it's the metadata system of record): consumer lag on the internal metadata consumer, under-replicated partitions. **Interpreting rising remote-copy lag — likely causes:** 1. Backend slow or throttling (object-store request limits, network), 2. RSM/credential/connectivity errors (check error metrics + logs), 3. Under-sized copier thread pool vs. partition count/ingest, 4. Ingest spike exceeding upload throughput, 5. Large backlog after enabling tiering on an existing topic (initial bulk upload). **Responses:** confirm via error metrics/logs; scale copier concurrency (remote.log.manager.thread.pool.size, copier-specific pool); raise backend throughput / request limits; throttle producers or temporarily reduce ingest; ensure local disk has buffer. For a fresh enablement, expect transient lag while the backlog uploads. **Edge cases:** copy lag that never drains after enabling tiering on a huge existing topic is normal initially but should trend down; if RemoteCopyErrorsPerSec is persistently non-zero, it's a config/permissions/backend problem, not just throughput. Watch that deletion (RemoteDelete) keeps pace or remote storage and request costs creep up.
- Why does rising RemoteCopyLag eventually threaten broker availability?A segment cannot be deleted from local disk until it's been uploaded. If copy lag grows, segments pile up locally even past local.retention, so local disk fills; a full log directory takes the broker offline.
- Copy throughput looks fine but RemoteCopyErrorsPerSec is non-zero. What does that tell you?It's not a throughput problem but a backend/config fault — bad credentials, missing bucket permissions, wrong endpoint/region, or transient object-store errors. Investigate logs and the RSM config rather than just scaling thread pools.
saying these in an interview costs you the question
- Only monitoring storage GB and ignoring the copy/delete pipeline and its lag.
- Treating all rising copy lag as a thread-pool issue while errors point to credentials/connectivity.
- Forgetting local disk fills when uploads stall (segments can't be deleted pre-upload).
- Ignoring remote-delete lag, which lets remote storage and request costs grow unbounded.