skip to content

Why does a database/sql service start up healthy when its database is down, and how do you fix readiness?

level: seniorimportance: should knowfreq 45%

answer

  1. startup never touched the server
  2. the handler proves the process, not the dependency
  3. make the check dial, under your own deadline
  4. restarting on a dependency blip makes it worse
  5. a scheduled run should exit non-zero

basics

~20 s

Because sql.Open never contacts the server, a process starts and reports ready with an unreachable database. Make the readiness check call db.PingContext with a short deadline, and keep liveness separate so a database blip does not restart the process.

solid answer

~50 s

The cause is that `sql.Open` is lazy: it builds a pool handle without dialling, so the only startup failure it can report is an unregistered driver name. A readiness endpoint that returns 200 because the process is running therefore says nothing about the database. Make readiness touch it: `db.PingContext(ctx)` under its own short deadline, returning 503 on failure. Keep liveness separate and database-free — if liveness pings, a brief outage restarts every replica and turns a dependency blip into your own. For a long-running server, prefer starting up, logging loudly and reporting not-ready until the ping succeeds, so the fleet heals when the database returns. For a process that starts and exits on a schedule there is no window to heal, so ping once at start and exit non-zero: a silent no-op run is worse than a visible failure.

code

go · 9 lines
go
func (s *Store) ready(w http.ResponseWriter, r *http.Request) {
	ctx, cancel := context.WithTimeout(r.Context(), 2*time.Second)
	defer cancel()
	if err := s.db.PingContext(ctx); err != nil {
		http.Error(w, "database unavailable", http.StatusServiceUnavailable)
		return
	}
	w.WriteHeader(http.StatusOK)
}

go deeper

for a junior

Remember why this happens: sql.Open does not connect, so a process can start with a dead database. Know that a readiness check has to ping or query for the answer to mean anything.

for a middle

Explain the mechanics of the fix: PingContext borrows a pooled connection and does a round trip, the deadline must be yours, and the handler returns 503 on failure rather than logging and continuing.

for a senior

Show the operational judgment: readiness versus liveness, why a database check in liveness causes restart storms, probe cost during an incident, and when to start not-ready versus exit non-zero.

for a principal

Own the convention across services: what every process must prove at boot, which dependencies are allowed to gate traffic, and how failure is surfaced so no scheduled job can quietly succeed while doing nothing.

## The mechanism behind the symptom `sql.Open` returns a `*sql.DB` without opening a connection. It resolves the driver name against the registry, possibly lets the driver pre-parse the DSN, and returns. No dial, no authentication. So a process configured with a dead host, a wrong password or a nonexistent database starts perfectly, and everything downstream that only asks "did startup return an error?" answers yes. If your health endpoint is the common `func(w http.ResponseWriter, r *http.Request) { w.WriteHeader(200) }`, it is telling the truth about the process and nothing about its dependencies. The orchestrator marks the instance ready, routes traffic to it, and users get the first honest error the system produces: a query failure inside a request. ## Make readiness prove the dependency ```go func (s *Store) ready(w http.ResponseWriter, r *http.Request) { ctx, cancel := context.WithTimeout(r.Context(), 2*time.Second) defer cancel() if err := s.db.PingContext(ctx); err != nil { http.Error(w, "database unavailable", http.StatusServiceUnavailable) return } w.WriteHeader(http.StatusOK) } ``` Three details matter here. **The deadline is yours.** Derive it from the incoming request's context so a cancelled probe stops work, but cap it yourself. Without a cap, a host that black-holes packets leaves the ping hanging and the probe's own timeout — set by someone else, in a manifest you may not own — becomes your failure semantics. A bounded ping fails in a time you chose. **A ping costs a pooled connection.** `PingContext` borrows a connection for the round trip, opening one if none is idle. A probe every second across many replicas is real, permanent load on the database's accept path, and during an incident it competes with the traffic you are trying to recover. If probes are frequent, cache the result for a short interval or run the check on a ticker and have the handler read the last outcome. **Ping proves less than a query.** It shows one connection could be established and is alive. It does not show that the schema is what you expect or that the account can read the tables you need. If those failures matter, a cheap real `SELECT` against a table you actually depend on is a stronger readiness signal — at the cost of a real query per probe. ## Liveness is a different question Liveness asks "is this process wedged and in need of a restart?" Readiness asks "should traffic come here right now?" Wiring a database check into liveness is one of the most reliable ways to convert a dependency outage into a self-inflicted one: the database wobbles, every replica fails liveness, every replica restarts, all of them come back cold and re-dial simultaneously, and the restart storm keeps the database from recovering. Liveness should test the process — that the request loop still schedules and responds. Readiness owns the dependency. ## Should startup fail, or should it start not-ready? This is the judgment call the question is really about, and the answer depends on process shape. **A long-running server.** Prefer to start. Log the failed ping loudly, expose it as not-ready, and keep retrying in the background. The instance then heals by itself the moment the database returns, with no deploy and no operator. Failing hard at boot instead means a database blip during a rollout leaves you with a crash-looping deployment, backoff delays, and possibly a rolled-back release that was fine. The nuance: start-up should still be *loud*. A service that starts, reports not-ready and logs nothing is the same silent failure in a different costume. **A short-lived process — a job that starts, works and exits, perhaps every minute.** The opposite. There is no traffic to withhold and no window in which healing helps: the run either does its work or it does not. Ping once immediately after `sql.Open` with a deadline well inside the schedule interval, and on failure log and exit non-zero so the scheduler records a failed run and your alerting sees it. The failure mode this prevents is the nastiest one on this whole topic: a job that starts every minute, silently processes nothing because its pool never connects, exits zero, and looks perfectly healthy on every dashboard while the data it was supposed to move quietly stops moving. A run that exits non-zero is a page; a run that exits zero having done nothing is discovered days later. ## What to check when you inherit the symptom Confirm the process really did no network work — a startup log with no connection line is the tell. Check whether the readiness handler references the database at all. Check whether the ping is bounded. Then check whether liveness is doing the database's job, because that determines whether your fix stops an outage or starts one.

  • Why should a database check stay out of the liveness probe?
    Liveness answers whether the process needs restarting. If a database outage fails liveness, every replica restarts at once, comes back cold and re-dials together, which prolongs the outage and can prevent recovery. Readiness withholds traffic without killing anything, which is the correct response to a dependency being down.
  • What does a probe that pings on every request cost you?
    Each ping borrows a connection from the pool and makes a round trip, opening a connection if none is idle. Across many replicas at a high probe frequency that is constant load on the database, worst exactly when it is already struggling. Cache the result briefly or run the check on a ticker and serve the last outcome.
  • Would a real query be a better readiness check than a ping?
    Sometimes. A ping proves only that a connection can be established and is alive. A cheap SELECT against a table you depend on also proves the account has privileges and the schema exists, catching failures a ping misses. The tradeoff is a real query per probe, so pick the cheapest statement that covers what you care about.
  • Should a service refuse to start when the database is unreachable?
    For a long-running server, usually not: start, log loudly, report not-ready and retry, so it heals when the database returns instead of crash-looping through a rollout. For a process that starts and exits on a schedule, do the opposite and exit non-zero, since there is no window in which healing helps and a silent zero-exit run hides the failure.

saying these in an interview costs you the question

  • Readiness returns 200 without touching the database
  • Puts the database ping in the liveness probe
  • Pings with no deadline, so the probe hangs
  • Assumes sql.Open failing is enough of a check
  • Lets a scheduled run exit zero having connected to nothing