skip to content

JUnit 5 offers assertLinesMatch for comparing two lists of strings. What can it express that a plain list-equality assertion cannot, and when is it the right tool?

level: seniorimportance: nice to knowfreq 20%

answer

  1. expected line = literal | full-match regex | >> marker >>
  2. >> 3 >> skips exactly 3; >> text >> skips many
  3. expected may not exceed actual line count
  4. for timestamps, ids, stack traces in text output
  5. parse JSON/XML instead of line-matching it

basics

~20 s

assertLinesMatch compares text line by line, but each expected line may be a literal, a regular expression matching the whole line, or a fast-forward marker like >> 3 >> or >> anything >> that skips lines. That lets you assert stable parts of multi-line output while tolerating timestamps, ids and unbounded filler.

solid answer

~50 s

`assertLinesMatch(List<String> expected, List<String> actual)` walks both lists in order and accepts an actual line if it is **equal** to the expected line, or if the expected line is a **regular expression** that matches the whole actual line, or if the expected line is a **fast-forward marker** — a line beginning and ending with `>>`, such as `>> 3 >>` to skip exactly three lines or `>> stack trace >>` to skip an unspecified number. That matters for output you cannot fully pin down: log or console output containing timestamps, generated ids, durations, host names, or a variable-length block you simply do not care about. With `assertEquals` on two lists you would have to either strip the volatile parts by hand or assert nothing at all. Use it for asserting rendered text — CLI output, generated files, templated messages. Keep the expected list anchored on the lines that carry meaning, and prefer a narrow regex over `.*` so the assertion still fails when the interesting part changes.

code

java · 13 lines
java
@Test
void consoleOutputMatches() {
    List<String> actual = runTool("--import", "orders.csv");

    assertLinesMatch(List.of(
        "katajob-cli 2.4.0",
        "Started at \\d{4}-\\d{2}-\\d{2}T.*",   // regex: timestamp varies
        ">> reading input >>",                   // skip any number of lines
        "Imported 120 orders, 0 rejected",       // the line that carries meaning
        ">> 2 >>",                               // skip exactly two lines
        "Done."
    ), actual);
}

go deeper

for a junior

Recognising that it exists and that expected lines may be regexes or skip markers is enough at this level.

for a middle

Be able to write the expected list correctly: full-match regex semantics, numeric versus free-text fast-forward markers, and the rule that expected cannot be longer than actual.

for a senior

Argue about assertion strength — which lines to pin, why every wildcard erodes the test, and when to parse structured output instead of line-matching it.

for a principal

Position it within a strategy for testing generated artefacts: golden-file approaches, normalising volatile fields at the source, and where approval testing beats inline expectations.

## The problem it solves Some outputs are multi-line text: a command-line tool's console output, a generated report or config file, a rendered email body, a formatted log. Asserting on them with `assertEquals(expectedList, actualList)` demands byte-perfect equality of every line, which fails the moment the output contains a timestamp, a UUID, an elapsed-time figure, a host name, or a stack trace whose length varies. The usual workarounds — regex-scrubbing the actual output before comparing, or asserting only `contains` on a couple of substrings — either hide bugs or make the test unreadable. `assertLinesMatch` in `org.junit.jupiter.api.Assertions` is JUnit 5's built-in answer: a positional, line-oriented matcher with two escape hatches built into the *expected* side. ## The three ways an expected line can match For each expected line, in order: 1. **Literal equality.** If the expected string equals the actual string, it matches. This is the common case and costs nothing. 2. **Regular expression.** Otherwise the expected line is treated as a regex and must match the *entire* actual line (a full match, not a find). So `Elapsed: \\d+ ms` matches `Elapsed: 42 ms`. Because literal equality is tried first, ordinary lines that happen to contain regex metacharacters usually still work — but a line containing, say, unbalanced brackets that is *not* literally equal will be attempted as a regex and may throw a pattern-syntax error, so escape when needed. 3. **Fast-forward marker.** An expected line that starts with `>>` and ends with `>>` and is at least four characters long is a marker rather than content. Its inner text controls the skip: if the text parses as an integer, exactly that many actual lines are skipped (`>> 3 >>`); otherwise the text is treated as a human-readable comment and an arbitrary number of actual lines is skipped until the next expected line matches (`>> stack trace >>`, `>> ... >>`). A trailing marker consumes the rest of the actual output. Structural rules: comparison is positional and in order, and the expected list may not contain more lines than the actual list. Failures report the line index and both texts, plus the remaining unmatched lines, which makes diagnosis much easier than a wall-of-text diff. Overloads accept `Stream<String>` on either or both sides and take an optional failure message or message supplier, which is useful when the same helper asserts several outputs. ## Where it earns its place - **CLI and console output.** Assert the banner, the important result lines and the exit summary, fast-forwarding over progress noise. - **Generated artefacts.** A generated source file, SQL script or configuration where a header carries a generation timestamp: match the header with a regex, the body literally. - **Formatted diagnostics.** Error output where a stack trace appears in the middle — `>> stack trace >>` skips it while still pinning the message before and the summary after. - **Template rendering.** Verify the fixed scaffolding of a rendered document while letting a couple of dynamic fields through as regexes. ## Where it is the wrong tool - **Structured data.** If the output is JSON, XML or YAML, parse it and assert on the parsed structure. Line matching on serialised structures makes the test sensitive to formatting, key order and whitespace, which is not the contract. - **Single strings.** For one line, a plain `assertEquals` or a regex assertion is clearer. - **Unordered output.** It is strictly positional; interleaved concurrent log output will be flaky. ## Keeping it honest The risk with a permissive matcher is a test that passes for everything. Two disciplines keep it useful: - **Prefer narrow regexes.** `Processed \\d+ records in \\d+ ms` still fails if the wording changes or the count line disappears; `.*` does not. Every `.*` you write is a line you have stopped testing. - **Fast-forward over noise only.** If you find yourself skipping most of the output, the assertion has stopped describing behaviour. Assert the lines that encode the outcome, and let the rest be skipped explicitly rather than by accident. Also remember locale and line separators: build the actual list with a locale-independent formatter and split on the platform separator (or normalise to `\n`) so the test does not depend on where it runs. ## Relationship to the other equality assertions `assertLinesMatch` is a specialisation, not a replacement. `assertIterableEquals` and `assertEquals` on lists remain the right choice when the text is genuinely deterministic — they are stricter and their intent is obvious. Reach for `assertLinesMatch` when the output has *deliberately* variable parts, and let the expected list document exactly which parts those are. That documentation value is the real payoff: a reader can see at a glance which lines are pinned and which are tolerated.

  • What is the difference between the fast-forward markers `>> 3 >>` and `>> anything >>`?
    A marker whose inner text parses as an integer skips exactly that many actual lines, so `>> 3 >>` consumes three lines and no more — the count is part of the assertion. A marker with non-numeric text is a human-readable comment and skips an arbitrary number of lines until the next expected line matches, so `>> stack trace >>` tolerates a block of unknown length. Use the numeric form when the count itself is meaningful and the free-text form when it genuinely is not.
  • Would you use assertLinesMatch to verify a JSON response body?
    No — JSON is structured data, and line matching makes the test depend on formatting, indentation and key ordering, none of which are part of the contract. Parse the body and assert on the resulting object or on specific paths, so a reformatted but semantically identical response still passes and a semantic change still fails. Line matching is for text whose line structure *is* the output, such as console or generated-file content.

saying these in an interview costs you the question

  • Thinking assertLinesMatch is just assertEquals for lists, with no regex or fast-forward capability.
  • Believing the regex only needs to be found somewhere in the line rather than matching the whole line.
  • Using .* for most expected lines, leaving an assertion that cannot fail.
  • Applying it to JSON or XML instead of parsing and asserting on structure.
  • Assuming it can match lines out of order or handle interleaved concurrent output.

context