skip to content

Compare three ways of pointing application traffic at whichever database node is currently the primary: a floating virtual IP, updating a DNS record, and a proxy layer such as HAProxy, PgBouncer or ProxySQL. What are the failure modes of each?

level: middleimportance: must knowfreq 50%

answer

  1. VIP = gratuitous ARP, needs L2, breaks live connections
  2. DNS = TTL + client caches, long tail of stale clients
  3. Proxy = per-connection decision, health checks, draining
  4. Proxy is a new SPOF → run it redundantly
  5. Driver multi-host = no hop, per-language logic

basics

~20 s

A floating IP moves in seconds but needs the nodes on one network segment and kills existing connections. DNS works anywhere but is hostage to TTL and client-side caching. A proxy decides per connection and can health-check and drain, but adds a hop and must itself be made highly available.

solid answer

~60 s

**Floating/virtual IP**: one address moves to the new primary, usually announced with gratuitous ARP. Clients need no config change and cutover is seconds, but it requires layer-2 adjacency (awkward in cloud, where it becomes an API call to reassign an address), existing TCP connections to the old holder break, and if the old node is not properly stopped two nodes can answer for the same address. **DNS**: change the record to the new primary. Works across networks and regions, needs no special infrastructure — but propagation is bounded by TTL plus every cache in between, and some runtimes cache resolved addresses effectively forever, so a subset of clients keeps talking to the demoted node. **Proxy**: clients always connect to the proxy; it decides per connection which backend is primary, health-checks continuously, can drain and briefly queue connections during promotion, and enables read/write splitting. Costs: an extra network hop and latency, another component to run, and it must be replicated or it becomes the single point of failure. Many stacks also use driver-level multi-host connection strings, which push the choice into the client.

code

text · 5 lines
text
# PostgreSQL JDBC: try both hosts, accept only a read-write session
jdbc:postgresql://db1:5432,db2:5432/appdb?targetServerType=primary

# read traffic: prefer a standby, fall back to the primary
jdbc:postgresql://db1:5432,db2:5432/appdb?targetServerType=preferSecondary&loadBalanceHosts=true

go deeper

for a junior

Name the three mechanisms and their headline tradeoff: fast but network-constrained, universal but slow to propagate, flexible but an extra component.

for a middle

Explain gratuitous ARP, TTL versus client-side caching, and why a proxy can drain connections while the other two cannot.

for a senior

Discuss the long tail of stale clients, making the proxy tier redundant, and how each mechanism behaves for connections that are already open.

for a principal

Frame it by topology and client control: how many networks, how many client languages, who owns the clients, and what the added hop costs against the operability the proxy buys.

## The problem After a failover the primary is a different machine. Something must tell clients where to send writes. There are three classic mechanisms, plus a fourth living in the driver. ## Floating (virtual) IP A service address is not bound to a machine; whichever node holds it answers for it. On promotion the new primary configures the address and broadcasts a gratuitous ARP so switches and peers update their tables. - **Good**: no client configuration at all; cutover in seconds; no extra process in the data path. - **Requires**: both nodes on the same layer-2 segment, or a cloud API that reassigns an elastic address — which turns "seconds" into "however long the control plane takes". - **Breaks**: every existing TCP connection to the old holder. Clients see resets or, worse, silently hung sockets if the old node vanished without sending anything, until keepalives or timeouts fire. - **Danger**: if the old primary is unhealthy but still running and still holding the address, two nodes can claim it. Preventing that is a fencing concern, not a routing one, but the routing layer is where it becomes visible. ## DNS Publish a name such as `db-primary.internal` and repoint it on failover. - **Good**: works across subnets, availability zones and regions; no special networking; trivially inspectable. - **Bad**: propagation. The record TTL is a floor, not a guarantee: resolvers, sidecars, connection pools and language runtimes each cache. JVM applications historically cache resolved addresses for the process lifetime unless the security property is tuned; some pool implementations resolve once at startup and never again. The result is a long tail of clients writing to a demoted node — which, if it has been correctly set read-only, gives errors rather than silent data divergence. - **Mitigation**: very low TTL, explicitly configured client-side DNS cache TTL, and pools that re-resolve on reconnect. This works, but it is a coordination problem across every client language you run. ## Proxy layer Clients connect to a stable proxy endpoint. The proxy holds backend connections, runs health checks, and routes each client connection (or each transaction) to the current primary. - **Good**: the routing decision is centralized and fast; the proxy can *drain* — stop handing out the old primary, let in-flight work finish, then cut. Some proxies can briefly pause and queue new connections during a switchover so clients see added latency instead of errors. It also enables read/write splitting, connection pooling, and per-client throttling, and gives one place to observe connection behaviour. - **Bad**: an extra hop (typically tens to hundreds of microseconds, plus a failure domain); its own configuration and upgrade lifecycle; and it is a new single point of failure unless it is run as a redundant pair or tier — which usually means putting a floating IP or a load balancer in front of the proxies, so you inherit one of the earlier mechanisms anyway, just at a layer where the state is cheap to move. ## Driver-level routing Modern drivers accept multiple hosts and a role predicate — connect to each in turn until one reports the required role. No extra infrastructure and no extra hop; but the logic and its rollout are per-language, discovery is only as fresh as the connection attempt, and there is no place to drain or queue. ## Choosing Same rack or same subnet, few clients, simple stack: a floating IP is hard to beat. Multi-zone or multi-region, or clients you do not control: DNS, accepting the tail, or a proxy in each zone. Heterogeneous fleets, read replicas, or a need to shape and observe connections: a proxy tier, made redundant. Most mature setups combine two — for example a redundant proxy pair fronted by a floating IP, with DNS naming the pair.

  • You lowered the DNS TTL to 5 seconds, yet after failover some application instances kept connecting to the old node for 20 minutes. Why?
    TTL only binds caches that honour it. Intermediate resolvers may clamp or ignore low TTLs, and clients frequently cache above DNS: many runtimes and connection pools resolve the hostname once and hold the address for the process lifetime or for a fixed cache interval. Fixing it means configuring the client-side resolver cache explicitly, forcing pools to re-resolve on reconnect, and validating with an actual failover drill rather than trusting the record.
  • If a proxy tier removes the single point of failure from the database, what removes it from the proxy?
    Redundancy plus a cheap routing mechanism in front of it. Proxies are stateless enough to run several instances, fronted by a floating IP pair, an anycast address, or a network load balancer with health checks; some teams run a proxy as a sidecar on every application host so its failure domain is one app instance. The design goal is that losing a proxy costs at most a reconnect, never a routing decision that has to be made by a human.

saying these in an interview costs you the question

  • Assuming a DNS TTL guarantees when clients switch
  • Thinking a floating IP migrates existing TCP connections rather than breaking them
  • Treating a single proxy instance as a highly available design
  • Believing routing alone prevents two nodes accepting writes
  • Assuming a virtual IP works unchanged in a cloud VPC across availability zones

context