A Testcontainers test times out waiting for its container only on CI. How do you diagnose it?
answer
- Get the container's own logs first
- Crash, wrong strategy, or genuinely slow
- Cold pulls and shared agents
- Timeout knob fixes only one of them
- Retrying start hides deterministic failures
basics
~20 sGet the container's own logs first, then separate the three causes: the container crashed, the wait strategy never matches what the image emits, or startup is genuinely slower on the agent. Only the last one is fixed by raising the startup timeout.
solid answer
~50 sStart from evidence rather than the timeout knob. Attach `withLogConsumer(new Slf4jLogConsumer(logger))` or capture `getLogs()` on failure so the container's stdout reaches the CI output — a `ContainerLaunchException` alone tells you nothing about *why*. Three shapes then separate cleanly. If the logs show an error and exit, the container is crash-looping: fix the image, environment or mounted config, and note that `withStartupAttempts` would only have retried the same crash. If the logs show the service coming up healthily and the wait still expired, the strategy does not match reality — a log regex that never matches the whole line, an HTTP probe on the wrong exposed port, or a healthcheck wait on an image without one. If the logs show a normal but slow start, that is the honest case for `withStartupTimeout(Duration.ofMinutes(2))`, plus attention to cold image pulls and agent contention. Being slower under load is not flakiness; a tight timeout tuned to a fast laptop is.
code
java · 12 linesGenericContainer<?> app = new GenericContainer<>("my/service:1.0")
.withExposedPorts(8080)
.withLogConsumer(new Slf4jLogConsumer(LoggerFactory.getLogger("container.app")))
.waitingFor(Wait.forHttp("/health").forPort(8080).forStatusCode(200))
.withStartupTimeout(Duration.ofMinutes(2));
try {
app.start();
} catch (ContainerLaunchException e) {
System.err.println(app.getLogs());
throw e;
}go deeper
Know that container output can be surfaced with a log consumer or getLogs(), and that a startup timeout is a symptom whose cause has to be read out of those logs.
Explain the three causes and their different fixes: crash, unobservable readiness, and genuine slowness. Know which knob addresses which, and why retrying start is not a general remedy.
Walk the whole diagnosis on a shared agent: evidence first, classification, the right fix, then the structural changes — image caching, pinned tags, job concurrency — that stop it recurring across the suite.
Own the reliability budget for container-backed CI: what the agents must provide, what timeout policy the codebase standardises on, and when a container-backed test is the wrong tool for the coverage it buys.
## Why "only on CI" is the important clue A developer machine and a build agent differ in ways that all bear on container startup: images are already cached locally but pulled cold on an ephemeral agent; the laptop is idle while the agent runs several jobs concurrently; the daemon may be local in one case and remote or nested in the other; CPU and memory limits may be far tighter. Timing-sensitive code that passes locally and fails on CI is almost always sensitive to one of those, not to the CI system being mysterious. ## Step one: make the container talk The default failure gives you an exception and a timeout, which is not a diagnosis. Two mechanisms bring the container's own output into the build log: - `withLogConsumer(new Slf4jLogConsumer(LoggerFactory.getLogger("container")))` streams container output into your test logging as it happens. - `container.getLogs()` retrieves the accumulated output, useful in a failure handler. Enable this permanently for container-backed suites; the storage cost is nothing and the alternative is guessing. Testcontainers also logs container startup activity itself, so raising its own log level adds pull and creation timing. ## Step two: classify the failure **The container never became healthy because it died.** The logs end in a stack trace, a missing environment variable, a permission error on a mounted file, or an out-of-memory kill. The wait strategy is a bystander here — it correctly reported that readiness never arrived. Fixing this is image and configuration work. Note especially that a container killed by the agent's memory limit looks identical to a crash, and that architecture mismatches (an amd64-only image on an arm64 agent, or the reverse) show up as immediate exits. **The container came up but the wait did not notice.** The logs show a normally started service; the strategy simply cannot observe it. Typical causes: a log regex written as a substring when the match is against the whole line; an HTTP probe pointed at the wrong exposed port or a path that returns 503 until warm; a healthcheck strategy on an image that declares no `HEALTHCHECK`. The tell is that the service is visibly ready in the logs well before the timeout expires. **The container is simply slow here.** Readiness arrives, just after the deadline. Cold image pulls on ephemeral agents, several jobs competing for CPU, and slower disks all contribute. This is where `withStartupTimeout` belongs, and where it is worth being generous — a two-minute ceiling costs nothing when startup takes eight seconds, and it removes a whole class of red builds. ## Step three: fix at the right level Raise the timeout when the cause is speed. Fix the strategy when the cause is observation. Fix the image or its configuration when the cause is a crash. Reach for `withStartupAttempts` only for genuinely intermittent infrastructure — it multiplies the time to fail and, applied to a deterministic problem, buries the evidence under repeated identical failures. Also consider structural fixes that remove the pressure entirely: pre-pulling or caching images on the agent so a cold pull is not on the critical path, pinning explicit image tags so the agent is not resolving a moving one, and reducing the number of containers a single job starts concurrently on a constrained machine. ## Distinguishing a timeout from a readiness lie These present differently, and the difference is worth saying out loud. A wrong-but-satisfiable strategy that returns *too early* does not time out at all: `start()` succeeds and the first query in the test fails with a connection reset or an empty schema. A strategy that never matches produces a startup timeout with a healthy container. Reading which of the two you have points straight at the fix. ## What a strong answer sounds like Evidence before knobs; three distinct causes named and separated; the timeout raised deliberately rather than reflexively; awareness that CI agents are cold, shared and constrained. A weak answer starts and ends with "increase the timeout and add retries", which fixes one of the three cases and hides the other two.
- How do you tell a crash-looping container from a wait strategy that never matches?Read the container output. A crash ends in an error and an exit before the deadline; a mismatched strategy shows a healthy, fully started service in the logs while the wait still expires. The two need completely different fixes, and only the logs distinguish them.
- When is withStartupAttempts the right tool?Only for genuinely intermittent infrastructure — an occasional registry or daemon hiccup. Against a deterministic failure it just repeats the same crash, multiplying the time to fail and burying the first, clearest error under identical retries.
- What structural changes reduce startup timeouts on CI without touching test code?Warm the image cache on the agent or pre-pull images, pin explicit tags so nothing is resolved at test time, give agents enough CPU and memory for the containers a job starts, and limit how many container-heavy jobs run concurrently on one machine.
- Why can a too-lenient wait strategy be worse than a timeout?Because it fails later and less legibly. start() succeeds, then the first query fails with a connection reset or a missing schema, and the failure looks like a test or application bug rather than a readiness problem. A timeout at least names the phase that failed.
saying these in an interview costs you the question
- Raises the timeout without reading container logs
- Adds retries to hide a crashing container
- Assumes CI flakiness rather than cold pulls and contention
- Tunes timeouts to a fast local machine
- Cannot distinguish a startup timeout from a too-early wait