skip to content

What are the practical limitations and operational risks of running row-level logical replication compared with block/WAL-level physical replication?

level: seniorimportance: must knowfreq 41%

answer

  1. DDL and sequences don't replicate — cutover must fix both
  2. Partial by design → never a promotable failover target
  3. Apply = real DML + index maintenance, limited parallelism
  4. Writable target → conflicts halt the entire subscription
  5. Unconsumed stream retains log on the SOURCE → disk outage

basics

~20 s

Logical replication does not carry schema changes, sequence values, or (by default) everything the cluster contains; apply is far slower per change; the writable target can hit conflicts that stall the whole subscription; the initial snapshot is expensive; and an unconsumed change stream retains log on the source until its disk fills.

solid answer

~1 min

Compared with physical replay, logical costs you: - **No DDL propagation.** Schema changes must be applied on both sides in the right order; a column that arrives before the subscriber has it stalls apply. - **Sequences and auto-increment counters are not replicated.** Their values must be advanced manually at cutover, or the target hands out duplicate keys. - **Incomplete by design.** Only published tables flow; roles, extensions, unpublished tables and often large objects do not. The target is therefore not a promotable copy of production. - **Much lower apply throughput.** Each event is real DML with index maintenance and constraint checks, generally single- or limited-parallel, and slower still if the target carries extra indexes. - **Conflicts stall everything.** The subscriber is writable, so a unique violation or a missing row typically halts the whole subscription until a human resolves or skips it. - **Initial snapshot cost.** Seeding large tables takes hours while changes accumulate and must be retained. - **Retention hazard.** A slow or dead subscriber pins change log on the source; that is a disk-exhaustion outage on production. Physical has none of these — it copies everything, cheaply, automatically — at the price of identical versions, all-or-nothing scope, and a read-only target.

go deeper

for a junior

Name the headline gaps — schema changes and sequences do not replicate, and the target only has the tables you published.

for a middle

Add apply cost, conflicts on a writable target, and the initial snapshot, and know the target is not a full copy of the cluster.

for a senior

Lead with the operational risks that page people: schema drift stalling apply, conflicts halting the whole subscription, and retained log filling the source's disk.

for a principal

Frame the estate-level position — physical for HA, logical for movement and integration — and specify the standing obligations: DDL choreography, permissions on replicated tables, retention alerting, and a re-seed threshold.

## The trade in one line Physical replication is *complete and cheap but rigid*. Logical replication is *flexible but partial and demanding*. Every limitation below is a direct consequence of the fact that logical replication reconstructs changes at the data level rather than copying storage, and therefore only knows about what it was told to carry. ## Limitation 1: schema changes do not travel A physical standby receives DDL automatically — a schema change is just more page-level records. Logical replication carries **data**, not catalog changes, so the subscriber's schema is the operator's responsibility forever. The practical failure is sharp: add a column on the publisher, and everything works until the first row event containing that column reaches a subscriber that lacks it. Apply fails, the subscription stops at that position, lag climbs, and log retention on the source starts growing. The discipline is a fixed ordering — additive changes on the subscriber first, then the publisher; drops in the opposite order; anything else coordinated behind a pause. This is a **permanent process cost**, and schema drift is the single most common cause of logical-replication incidents. ## Limitation 2: sequences and counters Row events carry the *values* a sequence produced, not the sequence's internal state. A subscriber therefore has the rows but a sequence still sitting at its initial value. Nothing is wrong until the subscriber becomes writable in earnest — at a cutover, say — at which point it starts allocating keys that already exist and immediately violates uniqueness. Every logical-replication cutover runbook must include advancing sequences past the maximum replicated value. ## Limitation 3: it is partial by construction Only published objects flow. What typically does *not*: unpublished tables, roles and permissions, extensions and their objects, large-object storage, and cluster-level settings. Some operations need explicit support to replicate at all (TRUNCATE is a classic case that needed dedicated handling rather than arriving free as a row event). The consequence that matters most: **a logical subscriber is not a promotable replacement for the source.** It cannot serve as the high-availability failover target, because it is missing objects an application depends on. If you need HA, you need a physical standby as well; logical does not substitute. ## Limitation 4: apply throughput Physical replay writes pages. Logical apply executes actual writes: locate the row, modify it, maintain every index on the target, evaluate constraints and triggers where configured, and write the target's own log. That is easily an order of magnitude more work per change, and apply parallelism is limited — often a single worker per subscription — while the publisher was written by hundreds of concurrent sessions. Two aggravating factors are common in practice: targets that carry *more* indexes than the source (reporting databases), and tables that lack a usable key so each event costs a scan. Both turn a throughput gap into a lag catastrophe under bulk writes. ## Limitation 5: conflicts on a writable target A physical standby cannot conflict — it is read-only. A logical subscriber is a normal database that accepts writes, so a locally inserted row can collide with a replicated one, or a locally deleted row can make an incoming UPDATE find nothing. The usual behaviour is that apply stops, and it stops for the **entire subscription**, not just the offending table. Resolution means a human deciding to fix the data or skip the transaction, and skipping means accepting silent divergence. Prevention is mostly access control: revoke write privileges on replicated tables from everyone except the replication role, and keep local objects in a separate schema. ## Limitation 6: the initial snapshot A subscription begins by copying the current contents of each published table, then streams changes from the snapshot point onward. For a large table this can run for hours, during which the publisher must retain everything that changed. This is the highest-risk window in the whole lifecycle: retention grows fastest exactly when the subscriber is least able to consume. Mitigations are per-table subscription staging, copying during low-write periods, and generous but monitored retention headroom. ## Limitation 7: retention as an outage vector This deserves separate billing because it is the one that takes production down. To guarantee no data loss, the publisher retains change log the subscriber has not consumed — tracked by a replication slot or its equivalent. A subscriber that is switched off, broken by schema drift, or simply too slow causes that retained log to accumulate **on the source's disk**. A full disk on a primary is a hard outage. So any logical setup needs: a monitored retained-log metric, an alert threshold well below the disk limit, and a documented decision point at which you drop the subscription and re-seed rather than protect a subscriber that is not coming back. The equivalent risk simply does not exist for a healthy physical standby with sane retention, and it is the reason "just add logical replication" is never free. ## What logical buys back Stating the limitations is only half an answer; a strong candidate closes with why anyone accepts them: selective replication, a writable and independently indexed target, cross-major-version streaming (and therefore near-zero-downtime upgrades), fan-in from multiple sources, and a change stream other systems can consume. Those capabilities are impossible physically. The mature position is that most estates run **physical for high availability and logical for movement and integration**, and budget the operational discipline logical demands.

  • Why is a logical subscriber usually unsuitable as a high-availability failover target?
    It replicates only published tables and carries none of the surrounding cluster state — roles, permissions, extensions, unpublished tables, sequence positions — so promoting it would give applications an incomplete database. Its apply throughput is also lower, so it typically trails further behind than a physical standby, which worsens the recovery point. Estates that need both capabilities run a physical standby for failover alongside the logical stream.
  • A subscription has stopped because of a unique-key conflict. What are the options, and what does each cost?
    You can fix the data on the subscriber so the pending transaction applies cleanly, which preserves consistency and is the preferred route. You can skip the offending transaction, which restarts apply immediately but leaves the target permanently divergent for that row. Or you can re-seed the affected table from the source, which is correct but pays the full snapshot cost and extends log retention on the publisher while it runs.
  • How would you monitor a logical replication setup so retention never takes down the source?
    Track the retained-log size per slot or subscription as a first-class metric on the source, alert well below the disk limit, and alert separately on subscription state — stopped, errored, or no longer connected — since those precede the growth. Define a hard threshold at which the runbook drops the subscription and re-seeds rather than continuing to retain, because protecting a dead subscriber is never worth an outage on production.

saying these in an interview costs you the question

  • Assuming DDL flows to a logical subscriber automatically
  • Forgetting to advance sequences at a logical cutover
  • Treating a logical subscriber as a disaster-recovery copy
  • Believing logical replication is faster because it transmits less data
  • Ignoring retained-log growth on the source when a subscriber is down
  • Resolving conflicts by routinely skipping transactions and calling the target consistent

context