Maintenance promoted the standby and the database was back within a minute, yet a nightly job kept failing for an hour - why?
answer
- the platform recovered, the client did not
- same endpoint name, different machine
- the pool holds what it holds
- no validation, no lifetime cap
- a process restart fixed it
basics
~20 sThe platform recovered; the client did not. Pooled connections opened before the promotion stayed in the pool, so every borrow handed the job a dead session, and a cached address plus no reconnect logic kept it failing until the process was restarted.
solid answer
~50 sA provider-initiated promotion severs every open connection. The service comes back in seconds behind the same endpoint name, but a client that holds a pool of long-lived connections does not notice: it keeps handing out sessions to a machine that is no longer the primary, and each borrow fails immediately. If the pool has no liveness check on borrow and no maximum connection lifetime, nothing evicts them. Add a resolved address cached in the process, a job with no reconnect path, and a transaction that rolled back mid-run with no resume point, and the hour of failures is the gap between the failover and someone restarting the job. The fixes are client-side: validate on borrow, cap connection lifetime, re-resolve the endpoint, retry the unit of work idempotently, and rehearse a failover instead of waiting for one.
code
pseudocode · 12 linesfunction borrowConnection(pool):
conn = pool.take()
if now() - conn.openedAt > maxConnectionLifetimeSeconds:
conn.close()
return pool.openNew() // re-resolves the endpoint name
if not conn.isAlive(): // cheap probe, runs before the caller gets it
conn.close()
return pool.openNew()
return conngo deeper
Take away the shape: a database failover ends every open connection, and an application that reuses pooled connections has to notice and open new ones before it works again.
Explain the mechanics - validation on borrow, a maximum connection lifetime, re-resolving the endpoint name - and why a promotion rolls back in-flight work rather than carrying it over.
Diagnose it as client-side: separate platform recovery time from application recovery time, name which pool setting kept the failure alive, and prove the fix by triggering a failover yourself.
Make restart tolerance a standard every service meets before it is allowed on a managed tier, so provider-initiated maintenance is absorbed by design rather than by whoever is on call.
## What actually happened on the platform side Maintenance on a replicated managed database usually takes the shortest available path: promote the standby replica to primary, then patch and rejoin the old primary as the new standby. From the platform's point of view this is a success story - the service answered again quickly, and the endpoint name the application uses still points at a working primary. Two details of that success are what the client trips over. First, **in-flight transactions did not survive**; they rolled back, because a promotion is a failover and not a restore of running work. Second, the endpoint name now resolves to a **different machine**, so every connection established before the promotion is attached to something that is no longer serving as primary. ## Why the outage outlived the failover - **The pool kept stale connections.** A connection pool exists to avoid the cost of opening sessions, so by design it holds them open. Nothing in the pool knows a failover happened. - **No validation on borrow.** If the pool hands out whatever it holds without a cheap liveness probe, the first thing the job learns is an error from a dead socket, once per attempt. - **No maximum connection lifetime.** With an unbounded lifetime, a connection that was opened at process start is still in the pool hours later; bounding the lifetime is what guarantees the pool eventually rebuilds itself even with no failure signal. - **A cached resolved address.** A process that resolved the endpoint name once and kept the result will keep dialling the old machine even when it does open a new connection. - **The unit of work was not resumable.** A long batch that aborted partway through had no checkpoint, so re-running it either was not safe or was not attempted. - **Nobody was watching for it.** The job logged connection errors, which look like noise, so an hour passed before a human restarted the process - and restarting the process is what finally fixed it, which is the giveaway that the fault was client-side. | Symptom | Underlying cause | Fix | |---|---|---| | Every query fails instantly after the failover | Pool handing out connections opened before promotion | Validate on borrow, evict on failure | | Failures persist far longer than the failover | Unbounded connection lifetime | Cap connection age so the pool refreshes itself | | New connections still go to the old machine | Resolved address cached in the process | Honour the record's expiry and re-resolve | | Batch cannot simply be re-run | No checkpoint and non-idempotent writes | Make the unit of work repeatable | ## Making the next promotion a log line 1. **Validate a connection before handing it out**, and evict rather than return a failed one. 2. **Bound connection lifetime**, so the pool rotates its sessions on a schedule of its own. 3. **Re-resolve the endpoint name** rather than pinning the address the process learned at startup. 4. **Retry the failed unit of work** a bounded number of times, and make it safe to repeat so a retried write does not double-apply. 5. **Fail fast on connect**, so a job discovers the new primary in seconds rather than blocking on a long timeout against a machine that will never answer. ## Proving it before the provider proves it for you The honest test is to cause the event yourself. Trigger a failover on a non-production copy, with the job running, and watch how long the client takes to recover on its own without a restart. That measurement is the only thing that distinguishes a system that tolerates provider-initiated restarts from one that has simply not had one recently. It is also the cheapest possible rehearsal: the promotion is exactly what maintenance will do, and a managed service gives you the button. The broader point is about where the boundary sits. Handing the database to a provider moves patching, host maintenance and failover mechanics off your plate, and in exchange the platform will restart your instance at moments you did not pick. The application's ability to reconnect is the part that never moved, and it is the part that decides whether a sixty-second promotion shows up as a metric or as an hour of failed work.
- Why does a long connect timeout make the recovery worse rather than safer?Because the connection the client is waiting on will never succeed. During a promotion the old machine stops answering as primary, so a generous timeout simply holds the job still while the new primary is already serving. A short connect timeout with a bounded retry finds the new primary in seconds.
- The failover took under a minute, so is this really a maintenance problem at all?The maintenance event is the trigger, not the fault. The platform did what it promises: promote and continue. The hour belongs to client behaviour that nobody exercised - a pool with no eviction, an address cached for the life of the process, a job with no resume point. Maintenance only made it visible.
- How do you know the fix works without waiting for the next maintenance window?Trigger a failover deliberately on a non-production copy with the workload running, and measure how long the client takes to serve normally again with no human intervention. If the number is seconds and no process was restarted, the client tolerates a promotion. If someone had to restart something, it does not.
saying these in an interview costs you the question
- Blames the provider because the failover itself was fast
- Assumes a connection pool notices a failover on its own
- Thinks the endpoint name changing is the problem, not the machine behind it
- Treats a process restart as the fix rather than the diagnosis
- Believes in-flight transactions are carried over to the promoted replica
- Raises the connect timeout to make reconnection more reliable