skip to content

How does schema linking work between two Schema Registries, and what role do exporters and IMPORT mode play?

level: seniorimportance: should knowfreq 25%

answer

  1. exporter = continuous schema replicator
  2. records carry schema ID → IDs must match on both sides
  3. destination context isolates subjects/IDs
  4. IMPORT mode → register explicit IDs/versions (subject must be empty)
  5. pairs with Cluster Linking for the data

basics

~20 s

Schema linking continuously replicates schemas from a source registry to a destination by configuring an exporter (a schema link). The exported schemas land in a context on the destination, which runs in IMPORT mode so it can register them with the source's exact IDs and versions, keeping IDs stable for replicated data.

solid answer

~50 s

Schema linking keeps subjects in sync across two Schema Registries — typically paired with Cluster Linking for the topic data. You create an **exporter** (also called a schema exporter / schema link) on the source (or via Confluent CLI / REST), specifying which subjects to replicate and the **destination context**. The exporter pushes schema versions to the destination continuously. Because consumers read the schema ID embedded in each record, the destination must preserve the source's IDs — so the target context is put in **IMPORT mode**, which permits registering schemas with explicit IDs and versions instead of auto-assigning them. Contexts isolate the replicated subjects/IDs from the destination's native ones. Exporters have a lifecycle (RUNNING/PAUSED) and you can monitor their status and offset. This underpins DR, migration, and global/aggregate registry topologies where data replicated by Cluster Linking must deserialize against identical schema IDs on the other side.

go deeper

for a junior

Know schema linking copies schemas from one registry to another.

for a middle

Know exporters do continuous replication into a context and that IDs need to be preserved.

for a senior

Explain IMPORT mode for explicit IDs, the empty-subject precondition, and pairing with Cluster Linking.

for a principal

Design DR/migration/aggregate topologies: context separation, compatibility handling during import, avoiding active-active loops.

## The core requirement: IDs must match Every Kafka record serialized with Schema Registry carries a **schema ID** (a magic byte plus a 4-byte ID). A consumer reads that ID and asks its registry for the schema. If you replicate *topic data* from cluster A to cluster B (via **Cluster Linking** or MirrorMaker), the records still carry A's schema IDs. If B's registry doesn't have those exact IDs mapped to the same schemas, deserialization breaks. So replicating data demands replicating schemas **with identical IDs**. ## Exporters (schema links) **Schema linking** solves this with an **exporter** — a configured, continuously-running replicator of schema versions from a source registry to a destination. You define: - the **subjects** to include (filter/patterns), - the **destination** registry connection, - the **destination context** the subjects land in, - subject renaming rules if needed. The exporter streams new and updated schema versions to the destination. Exporters are managed via the Confluent CLI (`confluent schema-registry exporter create/list/pause/resume/status`) or the REST API, and have a status (e.g. RUNNING, PAUSED) and a replication offset you can monitor. ## Contexts + IMPORT mode = ID preservation Two mechanisms make ID preservation possible: 1. **Context** — the replicated subjects go into a dedicated **context** on the destination so they don't collide with the destination's own subjects or its locally-assigned IDs. IDs are unique *within* a context. 2. **IMPORT mode** — normally the registry **auto-assigns** schema IDs and versions; you cannot dictate them. Putting a subject/context into **IMPORT mode** (a registry *mode*, alongside READWRITE/READONLY) allows registering schemas with **explicit IDs and versions**, so the destination reproduces the source's exact IDs. IMPORT mode requires the subject to be empty first (you can't import over existing versions). ## Typical topologies - **Disaster recovery / failover**: source → DR registry, so on failover consumers find identical schema IDs. - **Migration**: lift schemas to a new (e.g. Confluent Cloud) registry while preserving IDs. - **Aggregate/global registry**: many sources export into distinct contexts of one central registry. ## Import/export without continuous linking You can also do a **one-time export** of all subjects/versions (e.g. `confluent schema-registry ... export` or REST), then **import** them into a destination context in IMPORT mode — the manual equivalent of an exporter, used for migrations. ## Edge cases & gotchas - A subject/context must be **empty** to enter IMPORT mode and accept explicit IDs. - Compatibility checks may need adjusting during bulk import (often set to NONE during the import window, restored after). - Bidirectional/active-active linking needs careful context separation to avoid loops/collisions. - The exporter only handles **schemas**; the **data** is handled separately (Cluster Linking/MirrorMaker) — they must target compatible contexts/IDs. ## Why it matters Without schema linking, replicated topic data is undeserializable on the other cluster because the IDs don't resolve. Linking + contexts + IMPORT mode give you ID-stable, continuously-synced schemas for multi-cluster operations.

  • Why can't the destination just auto-assign new IDs to the replicated schemas?
    Replicated records still carry the source's schema IDs in their bytes. If the destination assigned different IDs, consumers reading those records couldn't resolve the embedded ID to the right schema and deserialization would fail.
  • What precondition must hold before a subject can enter IMPORT mode?
    The subject (in that context) must be empty — IMPORT mode lets you register explicit IDs/versions, which is only safe when there are no existing versions to conflict with.

saying these in an interview costs you the question

  • Saying schema linking also replicates the topic data — it replicates schemas; data needs Cluster Linking or MirrorMaker.
  • Claiming the destination assigns fresh IDs — the whole point is preserving source IDs via IMPORT mode.
  • Forgetting contexts — without an isolated context the imported IDs/subjects would collide with the destination's own.
  • Thinking you can enable IMPORT mode on a non-empty subject — it must be empty first.

context