skip to content

When would you run Debezium Server or the embedded engine instead of Debezium on Kafka Connect?

level: principalimportance: nice to knowfreq 30%

answer

  1. Same capture, three different hosts
  2. One of them has no broker at all
  3. Ask where the events are going
  4. Position storage moves with the shape
  5. One process means one point of failure

basics

~20 s

Debezium Server is a standalone process that streams change events to non-Kafka sinks; the embedded engine is a library running inside your own application. Choose them when there is no Kafka Connect cluster, accepting that you own restarts, scaling and offset storage.

solid answer

~50 s

Debezium ships in three shapes. **On Kafka Connect** it is a plugin in a managed cluster that handles connector lifecycle, offset storage and restarts for you — the default for anything Kafka-bound. **Debezium Server** is a standalone application that runs one source connector and writes to a configured sink such as Kinesis, Pub/Sub, Pulsar, Redis Streams or HTTP, chosen with `debezium.sink.type`; use it when your destination is not Kafka and standing up a Connect cluster would be pure overhead. **The embedded engine** is a Java library — you build a `DebeziumEngine`, hand it a consumer callback, and change events arrive in your own process; use it when the consumer is a single application that must react to changes without any broker in between. Both standalone shapes are single-process and store offsets in a file or similar local store, so durability of that store and process supervision become your problem, and there is no cluster to rebalance work onto.

code

properties · 9 lines
properties
debezium.sink.type=pubsub
debezium.sink.pubsub.project.id=analytics-prod

debezium.source.connector.class=io.debezium.connector.postgresql.PostgresConnector
debezium.source.topic.prefix=inventory
debezium.source.database.hostname=pg.internal
debezium.source.database.dbname=shop
-- must live on a durable volume, not container scratch
debezium.source.offset.storage.file.filename=/data/offsets.dat

go deeper

for a junior

Know that Debezium is usually run as a Kafka Connect plugin, and that a standalone server and an embeddable library also exist for cases where there is no Kafka.

for a middle

Explain what each shape hosts and where change events end up: a Connect cluster onto topics, a standalone server into a configured sink, the embedded engine into your own callback.

for a senior

Be concrete about what you take on without the framework — durable offset storage, process supervision, restart gaps, backpressure in your callback — and name the ephemeral-offset-store failure before it bites you.

for a principal

Own the selection criteria and their evolution: destination, source count, tolerable capture gap and team capacity, plus the migration path between shapes so the cheap starting point does not become a trap.

## Three deployment shapes Debezium's capture logic is the same everywhere; what differs is the runtime that hosts it and what happens to the events afterwards. **Debezium on Kafka Connect.** The connector runs as a plugin inside a Connect cluster. The framework owns connector lifecycle, task restarts, offset persistence in an internal topic, a management API and the ability to spread connectors across workers. Events land on Kafka topics. This is the default and the shape most interview answers assume. **Debezium Server.** A standalone application (built on Quarkus) that hosts exactly one source connector and forwards its events to a pluggable sink. Configuration lives in an `application.properties` file where source settings are namespaced under `debezium.source.*` and the destination is chosen with `debezium.sink.type` — Kinesis, Google Pub/Sub, Pulsar, Redis Streams, HTTP, RabbitMQ, NATS and others, plus Kafka itself. Offsets are stored wherever you point `debezium.source.offset.storage`, most simply a local file. **The embedded engine.** A library you add to your own JVM application. You construct a `DebeziumEngine`, supply the same connector properties you would give Connect, register a callback, and run it on an executor. Change events are delivered to your code; nothing is written anywhere unless you write it. Debezium Server itself is built on this engine. ## When each is the right call Reach for **Connect** when events are going to Kafka anyway, when several connectors need one operational model, or when you want the framework's restart and offset machinery rather than building it. The cost is running and upgrading a cluster. Reach for **Debezium Server** when the destination is not Kafka. If your architecture terminates in Pub/Sub or Kinesis, routing through Kafka purely to use Connect adds a system, a hop of latency and a second retention policy to reason about. Debezium Server removes that. It also suits small footprints — one database, one sink, one container. Reach for the **embedded engine** when the consumer is a single application and a broker would be ceremony. Typical uses: a service maintaining its own search index or cache from its database's changes, or a data-migration tool that needs the change stream for the duration of a cutover and then stops. It is also the only shape where you can apply arbitrary in-process logic to events before anything is persisted. ## What you give up Both standalone shapes are **single-process**. There is no cluster to move a connector onto when the host dies, no automatic restart of a failed task, no management API. You supply supervision — a container orchestrator restarting the pod is the usual answer — and you accept that during a restart, capture is stopped. **Offsets become your responsibility.** With Connect, position is persisted to an internal topic replicated across the cluster. With Debezium Server or the embedded engine, position typically lives in a local file. If that file is on ephemeral storage, a restart loses the position and the connector re-establishes from scratch — an expensive surprise on a large source, and a duplicate flood downstream. The offset store must be on a durable volume (or an equally durable alternative store), and it must be backed up and treated as state, not scratch. This is the single most common mistake with these shapes. **Scaling is bounded.** A Debezium Server instance hosts one source connector. Ten sources means ten instances, ten offset stores and ten sets of alerts. At that point the operational overhead you avoided by skipping Connect has quietly reappeared, distributed across ten deployments. **Delivery is at-least-once either way**, so sinks must be idempotent. The embedded engine additionally makes backpressure your problem: if your callback is slow, capture slows with it, and if your callback throws, you decide what happens next. Nothing retries for you. ## Making the decision A defensible framing weighs four things. *Destination* — Kafka or not — usually decides between Connect and Debezium Server on its own. *Number of sources* pushes toward Connect as the count grows, because per-instance overhead multiplies. *Availability requirement* asks how long a capture gap is acceptable: minutes of downtime during a pod restart is fine for a nightly analytics feed and unacceptable for a live cache. *Operational maturity* asks whether the team can genuinely run a Connect cluster, or whether one container with a persistent volume is the honest capacity. A reasonable answer also names the reversibility: because the connector configuration is nearly identical across shapes, starting embedded or on Debezium Server and later moving to Connect is a migration of the runtime and the offset store, not a rewrite of the capture. That makes the cheap option a defensible starting point rather than a lock-in.

  • What is the most common production mistake with Debezium Server?
    Leaving the offset store on ephemeral container storage. A restart then loses the recorded position, so the connector re-establishes from scratch: a full backfill of the source and a duplicate flood into the sink, at the worst possible moment. The offset file belongs on a durable volume, monitored and backed up like any other state.
  • You start with the embedded engine and later need three more consumers of the same change stream. What changes?
    The embedded model stops fitting — it delivers events to one process, so a second consumer means capturing the source twice, doubling load on the database. That is the point to introduce a broker: move to Connect or Debezium Server writing to a log every consumer can read independently. The connector configuration itself carries over largely unchanged.
  • Does running Debezium Server instead of Connect change the delivery guarantee?
    No. Capture is at-least-once in every shape, so sinks must be idempotent regardless. What changes is who provides the machinery around it: with Connect the framework persists position and restarts failed tasks, while a standalone instance leaves supervision and offset durability to your deployment.

saying these in an interview costs you the question

  • Assuming Debezium only runs inside a Kafka Connect cluster
  • Treating the offset file as scratch that can be recreated
  • Expecting a single Debezium Server instance to be highly available
  • Believing a non-Kafka shape provides stronger delivery guarantees
  • Running many single-source instances without counting the operational cost

context