skip to content

What is the RemoteStorageManager (RSM) plugin, and what do you need to configure to run one against an object store like S3 in production?

level: seniorimportance: should knowfreq 40%

answer

  1. RSM = data plane; RLMM = metadata plane
  2. interface only — bring a plugin (Aiven/Confluent)
  3. copyLogSegmentData / fetchLogSegment / deleteLogSegmentData
  4. config under impl.prefix: bucket, region/endpoint, creds, part/chunk size
  5. IAM role > static keys; monitor copy lag

basics

~20 s

The RSM is the pluggable component that physically reads/writes log segments and indexes to the remote backend. To run one against S3 you configure its class name plus backend settings: bucket, region/endpoint, credentials, chunk/part size, and optional encryption.

solid answer

~40 s

The RemoteStorageManager (RSM) is the KIP-405 plugin interface (org.apache.kafka.server.log.remote.storage.RemoteStorageManager) that handles the data plane of tiered storage: copyLogSegmentData, fetchLogSegment, fetchIndex, and deleteLogSegmentData. Kafka doesn't ship a production cloud backend, so you deploy a plugin (e.g. Aiven's tiered-storage-for-apache-kafka) and set remote.log.storage.manager.class.name to it. Plugin config lives under a prefix (e.g. remote.log.storage.manager.impl.prefix). For S3 you configure: the storage backend class, bucket name, region/endpoint (and path-style for S3-compatible stores like MinIO), credentials (IAM role/instance profile preferred over static keys), multipart upload part size, chunk size for ranged reads, optional client-side encryption (AES) and key management, and optional caching/compression. You also size the RLMM's __remote_log_metadata topic. Operationally you validate connectivity, set sane multipart sizes to bound memory, and monitor upload/fetch errors and latency.

go deeper

for a junior

Knows the RSM is the plugin that talks to the object store and must be configured with a bucket and credentials.

for a middle

Lists key configs (class name under prefix, bucket, region/endpoint, part/chunk size) and the RSM-vs-RLMM split.

for a senior

Reasons about credentials strategy, memory from part sizing, connectivity validation, and copy-lag monitoring.

for a principal

Owns plugin selection, security posture (least-privilege bucket policy, encryption), and capacity of the metadata topic across clusters.

**Where the RSM fits.** Tiered storage splits into two pluggable managers: the **RemoteLogMetadataManager (RLMM)** for the metadata plane (which segments exist remotely, their offsets, leader epochs), and the **RemoteStorageManager (RSM)** for the *data* plane — actually moving segment bytes and their index files to/from the backend. The RSM interface (`org.apache.kafka.server.log.remote.storage.RemoteStorageManager`) defines methods like `copyLogSegmentData` (upload a rolled segment + offset/time/transaction/leader-epoch indexes), `fetchLogSegment` / `fetchIndex` (ranged reads on the read path), and `deleteLogSegmentData` (when retention expires remotely). **Kafka ships an interface, not a cloud backend.** Apache Kafka provides the contract and a `LocalTieredStorage` reference impl for testing, but production users adopt a third-party RSM: Aiven's open-source `tiered-storage-for-apache-kafka`, Confluent's, or a vendor's. You point `remote.log.storage.manager.class.name` at the plugin's class. **Plugin configuration (prefixed).** RSM properties are namespaced under a configurable prefix, commonly `remote.log.storage.manager.impl.prefix` (and `remote.log.metadata.manager.impl.prefix` for the RLMM), so Kafka passes them through to the plugin. For an S3 backend you typically set: - **Backend selector / class** — which object-store driver the plugin uses (S3, GCS, Azure, HDFS). - **Bucket / container** name and key prefix. - **Region** and/or **endpoint URL** (custom endpoint + **path-style access** for S3-compatible stores such as MinIO/Ceph). - **Credentials** — prefer an **IAM role / instance profile / workload identity** over static access keys; the AWS SDK default credential chain is usual. - **Multipart upload part size** — bounds broker heap used while uploading large segments; too small ⇒ many parts, too large ⇒ memory pressure. - **Chunk size** — granularity for ranged GETs on the read path; affects fetch efficiency and the plugin's fetch cache. - **Encryption** — optional client-side AES encryption with a configured key (plus the object store's own SSE). - **Compression / caching** — some plugins compress chunks and maintain a local fetch cache. **Operational concerns:** - **Validate connectivity** before enabling topics: wrong region/endpoint/credentials shows up as upload failures and growing local disk (segments can't be deleted until uploaded). - **Memory:** part size × concurrent uploads bounds heap; tune `remote.log.manager.thread.pool.size` / copier concurrency. - **Metadata topic sizing:** `__remote_log_metadata` needs adequate partitions/replication and its own retention reasoning (it's the system of record for remote segments). - **Observability:** monitor RemoteCopyLag/RemoteCopyBytesPerSec, RemoteFetch errors, and per-broker upload/fetch latency; alert on rising copy lag (indicates backend or throughput problems). - **Security:** least-privilege bucket policy (PutObject/GetObject/DeleteObject/ListBucket on the prefix), encryption in transit (TLS endpoint) and at rest.

  • What's the difference between the RSM and the RLMM?
    RSM (RemoteStorageManager) is the data plane — it moves segment bytes and index files to/from the backend. RLMM (RemoteLogMetadataManager, default TopicBasedRemoteLogMetadataManager) is the metadata plane — it tracks which remote segments exist and their offset/epoch mapping, stored in __remote_log_metadata.
  • Why prefer an IAM role over static access keys for the RSM, and how does the multipart part size affect brokers?
    IAM roles/instance profiles avoid long-lived secrets in broker config (rotation, least privilege, no leakage). Multipart part size bounds broker heap during uploads: larger parts mean fewer requests but more memory per concurrent upload, so it's tuned against thread-pool concurrency.

saying these in an interview costs you the question

  • Claiming Apache Kafka ships a built-in production S3 RSM (it ships only an interface + a local test impl).
  • Confusing the RSM (data) with the RLMM (metadata).
  • Hardcoding static AWS keys in server.properties instead of using a role/credential chain.
  • Ignoring that failed uploads stall local-segment deletion and grow disk.

context