In a Docker Compose file, does the `depends_on` key guarantee that a dependency such as a database is ready to accept connections before the dependent service starts? How do you make startup ordering actually reliable?
answer
- depends_on = start order, not readiness
- healthcheck + condition: service_healthy
- start_period = grace window for slow boots
- service_completed_successfully = migrations
- apps retry anyway; ordering is not a guarantee
basics
~20 sNo. Plain depends_on only orders container start, not readiness — the database process may still be initialising. Add a healthcheck to the dependency and use depends_on: <svc>: condition: service_healthy, and still make the app retry its connection.
solid answer
~50 s`depends_on` in its short list form only controls **start order and shutdown order**: Compose starts the dependency's container first. The container being started says nothing about the process inside having finished initialising, so an app that connects on boot will still hit "connection refused" against a Postgres that is mid-initdb. The fix has two halves: 1. **Declare readiness.** Give the dependency a `healthcheck` (a command like `pg_isready` with `interval`, `timeout`, `retries`, `start_period`), then use the long form: `depends_on: { db: { condition: service_healthy } }`. Compose then blocks the dependent container's start until the healthcheck passes. `service_completed_successfully` covers one-shot jobs like migrations. 2. **Do not rely on it alone.** Healthchecks help at startup only; a dependency can restart later, and in a cluster nothing orders you at all. Applications should retry with backoff on connection failure. `docker compose up --wait` blocks the CLI until healthy, which is what CI wants.
code
yaml · 27 linesservices:
db:
image: postgres:16
environment:
POSTGRES_PASSWORD: devpassword
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]
interval: 5s
timeout: 3s
retries: 10
start_period: 20s
migrate:
image: myapp/migrations:latest
depends_on:
db:
condition: service_healthy
api:
image: myapp/api:latest
ports:
- "8080:8080"
depends_on:
db:
condition: service_healthy
migrate:
condition: service_completed_successfullygo deeper
Know that depends_on only orders container start and that a database may not be ready yet; recognise the connection refused on first up.
Write the healthcheck plus long-form depends_on condition correctly and explain interval/retries/start_period.
Lead with 'ordering is not a readiness guarantee', require application-side retry with backoff, and use up --wait in CI with a bounded timeout.
Frame it as a dependency-failure model: startup ordering is a local convenience, while steady-state correctness comes from retries, circuit breaking and readiness signalling that hold wherever the image runs.
## What `depends_on` actually does In its short form: ```yaml depends_on: - db ``` `depends_on` tells Compose two things: start `db`'s container **before** the dependent service's container, and stop it **after** on `down`. That is the whole contract. "Container started" means the container process has been created and the entrypoint has been exec'd — nothing more. A Postgres image at that moment may be running `initdb`, creating the role, and restarting internally; a JVM app may be seconds away from binding its port. The classic symptom is a dependent service that crashes on first `up` with `connection refused` or `could not translate host name`, then works after `docker compose up` is run a second time (because by then the database is warm). Candidates who "fix" this by adding `sleep 10` to the entrypoint have papered over it: the sleep is both too long on a fast machine and too short on a cold CI runner. ## Declaring readiness with `healthcheck` A container has a health state (`starting`, `healthy`, `unhealthy`) only if a healthcheck is defined — either baked into the image (`HEALTHCHECK` in the Dockerfile) or declared in the compose file: ```yaml healthcheck: test: ["CMD-SHELL", "pg_isready -U postgres -d app"] interval: 5s timeout: 3s retries: 10 start_period: 30s ``` The fields matter: - **`test`** — the command run *inside the container*. `CMD` takes an exec-form argv; `CMD-SHELL` runs a shell string. Exit code 0 = healthy. The binary must actually exist in that image — `curl` is absent from many slim images, which is why a healthcheck can be permanently unhealthy for a service that works fine. - **`interval`** — how often to probe once running. - **`timeout`** — how long a single probe may take before it counts as a failure. - **`retries`** — consecutive failures before the state flips to `unhealthy`. - **`start_period`** — a grace window during which failures do not count toward `retries`, for slow-booting services. During this window the state is `starting`. A good healthcheck tests the thing dependents need, not merely that a process exists. `pg_isready` for Postgres, `redis-cli ping` for Redis, an HTTP `/health` endpoint for an app. Checking only that a TCP port is open is weaker but still far better than nothing. ## Wiring readiness into ordering The long form of `depends_on` accepts a condition per dependency: ```yaml depends_on: db: condition: service_healthy migrate: condition: service_completed_successfully cache: condition: service_started ``` - **`service_started`** — the old behaviour: start order only. - **`service_healthy`** — block until the dependency's healthcheck reports healthy. If it never does, Compose fails the `up` with a clear error instead of leaving you to read logs. - **`service_completed_successfully`** — block until the dependency's container exits with code 0. This is how one-shot jobs are sequenced: a `migrate` service that runs schema migrations and exits, gating the API. Adding `restart: true` under a dependency (supported in recent Compose) makes Compose restart the dependent when the dependency is restarted. ## Why applications must still retry Ordering at startup is a convenience, not a correctness guarantee: - A dependency can crash and restart at any time after boot; nothing re-orders anything then. - Healthchecks are polled, so there is a window where a service is reported healthy but is shutting down. - The same image will eventually run somewhere with no ordering primitive at all, where every pod starts concurrently. So the durable answer is: connect with retry and bounded backoff, fail the readiness probe while unconnected, and treat compose ordering as a way to reduce noise during local `up`, not as the mechanism that makes the system correct. Interviewers listen specifically for that sentence. ## CI ergonomics `docker compose up -d --wait` returns only when all services with healthchecks are healthy (and fails if any becomes unhealthy), which removes the polling loop most CI scripts hand-roll. `--wait-timeout` bounds it. Without healthchecks, `--wait` has nothing to wait on — another reason to define them even for services nothing depends on.
- The healthcheck for a service is defined but the container is stuck in `starting` forever. How do you debug it?Run `docker inspect --format '{{json .State.Health}}' <container>` to see the last probe outputs and exit codes, which Compose does not print. The usual causes are a binary that does not exist in a slim image (curl/wget missing), a probe pointing at localhost when the process binds only an external interface, or a timeout shorter than the real response time. Reproduce by running the exact `test` command with `docker compose exec`.
- Why is adding `sleep 15` before the app starts a bad substitute?It is a guess that is simultaneously wasteful and unreliable: too long on a warm laptop, too short on a cold CI runner or a large database restore. It also does nothing about the dependency restarting later. A healthcheck expresses the actual condition, and application-side retry handles the general case.
- Does `depends_on` affect anything besides startup?Yes — it also defines shutdown order (dependents are stopped before their dependencies on `down`), and `docker compose up <service>` implicitly starts that service's dependencies. Recent Compose also supports `restart: true` under a dependency so the dependent is restarted when the dependency is.
depends_on is like being told the kettle was switched on before you were called to the kitchen; the healthcheck condition is waiting until it actually whistles.
saying these in an interview costs you the question
- Believing plain `depends_on` waits for the dependency to be ready to serve traffic.
- Using a fixed `sleep` in an entrypoint instead of a healthcheck plus retry.
- Writing a healthcheck that calls `curl` in an image that has no curl, then concluding healthchecks are broken.
- Assuming healthcheck-gated ordering removes the need for connection retry logic in the application.
- Confusing `service_completed_successfully` with `service_healthy` for a one-shot migration container that exits.