What are the RemoteStorageManager (RSM) and RemoteLogMetadataManager (RLMM) plugins, and how do their responsibilities differ?
answer
- RSM = bytes / data plane
- RLMM = metadata / control plane
- RemoteLogSegmentMetadata: offsets + epoch + state
- default RLMM = __remote_log_metadata topic
- COPY_STARTED -> COPY_FINISHED, DELETE_STARTED -> DELETE_FINISHED
basics
~10 sRemoteStorageManager moves the actual segment bytes to/from remote storage (the data plane). RemoteLogMetadataManager tracks metadata — which segments exist remotely, their offset/epoch ranges, and their state (the control plane).
solid answer
~40 sKIP-405 defines two pluggable Java interfaces. **RemoteStorageManager (RSM)** is the data-plane plugin: it knows how to copy a closed log segment (plus its indexes — offset, time, transaction, producer-snapshot, leader-epoch) to the object store, fetch byte ranges back during reads, and delete remote segments. It abstracts S3/GCS/HDFS specifics. **RemoteLogMetadataManager (RLMM)** is the control-plane plugin: it stores and serves **RemoteLogSegmentMetadata** — segment ID, start/end offsets, leader-epoch lineage, timestamps, and lifecycle state (COPY_SEGMENT_STARTED, COPY_SEGMENT_FINISHED, DELETE_SEGMENT_STARTED, DELETE_SEGMENT_FINISHED). It answers 'which remote segment holds offset X?'. The default RLMM, TopicBasedRemoteLogMetadataManager, persists this metadata in an internal Kafka topic, __remote_log_metadata. Brokers' RemoteLogManager orchestrates both. Separating data from metadata lets you swap object-store backends without changing how Kafka tracks segment placement.
go deeper
Remember the split: one plugin moves bytes (RSM), the other remembers what was moved (RLMM).
Name the interface methods (copyLogSegmentData, fetchLogSegment, deleteLogSegmentData) and the metadata states, and the default __remote_log_metadata topic.
Explain why metadata lives in a Kafka topic, how the STARTED/FINISHED states give crash safety, and how RLM orchestrates both plugins.
Discuss tradeoffs of the topic-based metadata store (ordering, partition mapping, scale), pluggability for multi-cloud, and durability invariants RSM must uphold before FINISHED.
## Why two plugins Tiered storage has two fundamentally different jobs: 1. **Move and retrieve bytes** to/from an external store (which could be S3, GCS, Azure Blob, HDFS, or a vendor system). 2. **Remember what was moved** — for each remote segment, which offsets it covers, what leader epochs it spans, when it was written, and what lifecycle state it's in — so that a read for offset X can be routed to the right remote object. KIP-405 deliberately splits these into two interfaces so each can be implemented and scaled independently. ## RemoteStorageManager (RSM) — the data plane Interface: **org.apache.kafka.server.log.remote.storage.RemoteStorageManager**. Core methods: - **copyLogSegmentData(metadata, segmentData)** — uploads a closed segment file plus its associated index files: the **offset index**, **time index**, **transaction index**, **producer-snapshot**, and the **leader-epoch checkpoint**. These indexes are needed so reads can locate positions without the broker holding the data locally. - **fetchLogSegment(metadata, startPosition[, endPosition])** — returns an InputStream over a byte range of the remote segment (this is what serves a remote read). - **fetchIndex(metadata, indexType)** — retrieves a specific index for a remote segment. - **deleteLogSegmentData(metadata)** — removes the remote object when retention expires. The RSM hides all object-store specifics (auth, bucket layout, multipart upload, retries). ## RemoteLogMetadataManager (RLMM) — the control plane Interface: **org.apache.kafka.server.log.remote.storage.RemoteLogMetadataManager**. It manages **RemoteLogSegmentMetadata** records, each carrying: - a unique **RemoteLogSegmentId**, - **startOffset / endOffset**, - **leader-epoch → start-offset** map (segment epoch lineage), - **maxTimestamp**, broker/event timestamps, - a **state** in a small state machine: COPY_SEGMENT_STARTED → COPY_SEGMENT_FINISHED, and later DELETE_SEGMENT_STARTED → DELETE_SEGMENT_FINISHED. Key lookups it serves: - **remoteLogSegmentMetadata(partition, leaderEpoch, offset)** — 'which remote segment contains this offset for this epoch?' (drives the read path). - **highestOffsetForEpoch / listRemoteLogSegments** — used for retention and recovery. The default implementation, **TopicBasedRemoteLogMetadataManager**, stores these records in an internal compacted-style Kafka topic called **__remote_log_metadata**, partitioned so each user partition's metadata maps to a metadata partition. This makes the metadata itself replicated and durable using Kafka's own machinery. ## How they fit together On each broker, the **RemoteLogManager (RLM)** component runs background tasks per leader partition. It (1) detects closed segments eligible to offload, (2) writes a COPY_SEGMENT_STARTED metadata record via RLMM, (3) calls RSM.copyLogSegmentData, (4) writes COPY_SEGMENT_FINISHED. On reads it asks RLMM for the segment covering the requested offset, then streams bytes via RSM.fetchLogSegment. For retention it marks DELETE_SEGMENT_STARTED, calls RSM.deleteLogSegmentData, then DELETE_SEGMENT_FINISHED. ## Edge cases - The two-phase STARTED/FINISHED states make offload and delete **idempotent and crash-safe**: if a broker dies mid-copy, a segment left in COPY_SEGMENT_STARTED is not yet readable and can be retried. - Metadata correctness is critical: a segment in COPY_SEGMENT_FINISHED but whose bytes are missing remotely would break reads, so RSM must guarantee durability before FINISHED is written. - Both plugins are configured via classpath + config (remote.log.storage.manager.class.name, remote.log.metadata.manager.class.name).
- Where does the default RemoteLogMetadataManager store its metadata?In an internal Kafka topic named __remote_log_metadata, managed by TopicBasedRemoteLogMetadataManager — so the metadata is itself replicated and durable via Kafka.
- Why are segment states split into COPY_SEGMENT_STARTED and COPY_SEGMENT_FINISHED?For crash-safe, idempotent offloads: a segment only becomes readable after FINISHED, so an interrupted copy (still STARTED) can be safely retried without serving incomplete data.
saying these in an interview costs you the question
- Saying the RSM tracks offsets/metadata — that is the RLMM's job.
- Saying the RLMM uploads segment bytes — that is the RSM's job.
- Claiming metadata is stored in ZooKeeper — the default RLMM uses the __remote_log_metadata Kafka topic.
- Forgetting that indexes (offset/time/txn/leader-epoch) are uploaded alongside the segment, not just the data file.