skip to content

In Testcontainers, how do you capture a container's logs when start() fails and the test dies?

level: seniorimportance: should knowfreq 34%

answer

  1. Evidence must be captured while it is alive
  2. One line per container, at build time
  3. Streamed frames survive the failure
  4. start() throws a launch exception
  5. Snapshot access needs a container that still exists

basics

~20 s

Attach a log consumer such as Slf4jLogConsumer before start() so container output streams into the test log as it happens, and read the ContainerLaunchException that start() throws. For a container that did start, getLogs() returns its output as a snapshot.

solid answer

~50 s

The failure you get by default is unhelpful: `start()` throws `ContainerLaunchException` after the readiness check gives up, and the container itself may already be gone, so `getLogs()` has nothing to read. The fix is to stream rather than to inspect after the fact. Attach `withLogConsumer(new Slf4jLogConsumer(logger))` when you build the container, so every stdout and stderr frame lands in your test log with the rest of the run — in CI that is the only artefact you will have. Then read the exception message, which names the readiness check that timed out, and compare it against what the streamed output shows: a process that crashed on bad configuration looks completely different from one that is healthy but slower than the startup timeout. Turning up logging for Testcontainers' own logger adds the pull, create and start steps around it.

code

java · 8 lines
java
private static final Logger log = LoggerFactory.getLogger(PaymentsIT.class);

GenericContainer<?> app = new GenericContainer<>(DockerImageName.parse("acme/payments:1.4.2"))
        .withExposedPorts(8080)
        .withLogConsumer(new Slf4jLogConsumer(log).withPrefix("payments"))
        .withStartupTimeout(Duration.ofSeconds(90));

app.start();   // on failure: ContainerLaunchException, but the streamed output is already in the log

go deeper

for a junior

Know that container output is not printed by default and that a log consumer must be attached before starting the container if you want to see it.

for a middle

Explain why a snapshot accessor is useless after a failed start while a streaming consumer is not, and what the launch exception does and does not tell you.

for a senior

Demonstrate the triage: crashed process, slow-but-healthy process, wrong readiness check, or image never pulled — and pick a different fix for each rather than reaching for a longer timeout.

for a principal

Make diagnosability a standard: log consumers and framework logging enabled by convention in the shared test base, so a red CI job is self-explaining instead of needing local reproduction.

## Why the default failure is opaque When a container never becomes ready, `start()` blocks until the configured timeout and then throws `ContainerLaunchException`. The message tells you that the readiness check did not pass — it does not tell you *why*, because Testcontainers is outside the container looking in. Worse, by the time you are holding the exception, the failed container has usually been removed, so any attempt to inspect it after the fact finds nothing. Locally you can sometimes race `docker logs`; in CI you cannot. So the discipline is: arrange for the evidence to be captured *while* the container is alive, before you need it. ## Stream the output with a log consumer `withLogConsumer(Consumer<OutputFrame>)` attaches a sink that receives every stdout and stderr frame as the container produces it. The built-in `Slf4jLogConsumer` forwards those frames to an SLF4J logger, and `withPrefix("…")` labels them so a multi-container suite stays readable. Because frames are delivered as they arrive, output produced by a container that dies two seconds in is already in your log by the time the launch exception is thrown. This is the single highest-value habit for containerised tests, and it costs one line per container. On a build agent it turns "the job failed and I have no idea why" into "the database refused to start because the data directory was not empty", which is usually right there in the first ten lines of the image's own output. ## getLogs() and its limits `getLogs()` returns the container's accumulated output as a `String` snapshot, with overloads to select stdout or stderr. It is excellent for assertions and for dumping state at the end of a *successful* start — for example, asserting that a broker announced the address you expected. It is a poor debugging tool for a failed start, precisely because it needs a container that still exists. Treat the consumer as the diagnostic path and `getLogs()` as the assertion path. ## Read the exception, then classify the failure With the streamed output in hand, launch failures fall into a few recognisable classes: - **The process crashed.** The output ends in a stack trace or an image-specific error, and the container exited. Nothing about timeouts will help; fix the configuration you passed in. - **The process is healthy but slow.** The output shows normal progress that simply had not finished — a database initialising a fresh cluster, a JVM-based service warming up — and the run is on a loaded or cold agent. Here `withStartupTimeout(Duration)` is the honest fix. - **The process is up but the readiness check is looking at the wrong thing.** Output says "ready to accept connections" yet the check still times out: the check is watching a port the process does not bind, or a path that returns a non-success status. - **Nothing started at all.** The failure is upstream of the container — the image could not be pulled, or the architecture does not match the host, and the output is empty. That is a different investigation entirely. ## Turn up Testcontainers' own logging Testcontainers logs its own steps — resolving the image, pulling, creating, starting, and each readiness attempt — under its own logger namespace. Raising that logger to DEBUG in the test logging configuration surrounds the container's output with the framework's view of the timeline, which is how you tell a slow *pull* apart from a slow *start*. That distinction matters on CI agents with cold image caches, because the fix is caching, not a longer startup timeout. ## Retries are a diagnosis, not a cure `withStartupAttempts(int)` retries the whole create-and-start cycle. It is defensible for a genuinely flaky external dependency and indefensible as a way of muting a real failure: it multiplies the time a broken test takes to fail and hides the frequency of the underlying problem. If you reach for it, record why, and keep the log consumer attached so the failed attempts remain visible. ## What a strong answer sounds like Name the exception, name the consumer, and — most importantly — show the reasoning that separates "crashed" from "slow" from "the readiness check is wrong", because that triage is the actual senior skill. A weak answer proposes raising the timeout as the first move, which fixes one of the four cases and wastes minutes per run in the other three.

  • How do you tell a slow image pull apart from a slow container start?
    Raise Testcontainers' own logger to DEBUG: it logs image resolution and pulling separately from creating and starting the container, so the timeline shows which phase consumed the time. The fixes differ — a slow pull is solved by warming or mirroring the image cache on the agent, a slow start by a longer startup timeout.
  • When is increasing the startup timeout the right fix rather than a cover-up?
    When the streamed output shows the service progressing normally and simply not finishing in time — a fresh database cluster initialising, or a cold, loaded CI agent. If the output ends in a crash or the service reports itself ready while the check still fails, more time changes nothing and only slows the failure down.
  • Why is withStartupAttempts risky as a default?
    It converts a deterministic failure into a slower, noisier one and hides how often the underlying problem occurs. Every retry pays the full create-and-start cost, so a broken container turns a fast red into a multi-minute red, and the flakiness never gets measured or fixed.

saying these in an interview costs you the question

  • Raises the startup timeout before reading any logs
  • Calls getLogs() after the container was already removed
  • Adds startup retries to silence a real crash
  • Assumes the exception message explains the cause
  • Debugs by hand locally with no CI artefact

context