How does JUnit 5's @CsvFileSource load its data, what is the difference between its resources and files attributes, and what is numLinesToSkip for?
answer
- resources = classpath (leading / = root, else package-relative)
- files = filesystem, resolved against working directory
- numLinesToSkip = 1 for a header; default 0
- Default quote char here is " (inline @CsvSource uses ')
- encoding UTF-8, lineSeparator \n, maxCharsPerColumn guard
basics
~20 s@CsvFileSource reads rows from CSV files instead of the annotation. resources names classpath resources (usually under src/test/resources); files names filesystem paths. numLinesToSkip skips the leading lines — set it to 1 to skip a header row. It defaults to 0.
solid answer
~50 s`@CsvFileSource` is `@CsvSource` with the table moved into real `.csv` files. Each non-skipped line is one invocation; columns bind positionally exactly as inline. - **`resources`** — classpath resources. A leading slash makes the path absolute from the classpath root (`/data/users.csv`); without it, the path resolves relative to the test class's package. This is the portable choice: the file ships in the test JAR and works on any machine and in CI. - **`files`** — plain filesystem paths, resolved against the working directory. Only reach for it when the data genuinely lives outside the build, because the working directory differs between IDE and build tool. You may list several files; their rows are concatenated. `numLinesToSkip` drops leading lines — `numLinesToSkip = 1` for a header row; default `0`. Other knobs: `delimiter`/`delimiterString`, `quoteCharacter` (default `"` here, unlike inline `@CsvSource`), `nullValues`, `encoding` (default UTF-8), `lineSeparator` (default `\n`), and `maxCharsPerColumn`.
code
java · 5 lines@ParameterizedTest
@CsvFileSource(resources = "/tax-brackets.csv", numLinesToSkip = 1, nullValues = "N/A")
void computesTax(int income, String region, BigDecimal expected) {
assertEquals(expected, TaxCalculator.of(region).tax(income));
}go deeper
Know that the data lives in a .csv on the test classpath, that resources points at it, and that numLinesToSkip = 1 skips the header.
Distinguish resources from files and their path-resolution rules, and list the other knobs: delimiter, quoteCharacter, nullValues, encoding.
Reason about portability — classpath resources for anything that ships, filesystem paths only for externally supplied data — and recognise encoding and maxCharsPerColumn failures from their symptoms.
Decide who owns the fixture: externalising the table is worth it when non-developers edit it or several tests share it, and a cost when it splits the reviewer's attention across two files.
## What it is `@CsvFileSource` supplies arguments to a `@ParameterizedTest` from one or more CSV files. Semantically it is identical to `@CsvSource` — each data line becomes one invocation, columns are split on a delimiter and bound positionally to the test method's parameters, and each column is converted from text to the declared parameter type. The only difference is *where the table lives*. ```java @ParameterizedTest @CsvFileSource(resources = "/tax-brackets.csv", numLinesToSkip = 1) void computesTax(int income, String region, BigDecimal expected) { assertEquals(expected, TaxCalculator.of(region).tax(income)); } ``` ## resources versus files These are two different lookup mechanisms and choosing wrongly is the classic cause of "works in my IDE, fails in CI". **`resources`** goes through the classpath. Put the file under `src/test/resources` and the build tool copies it into the test classpath; the annotation then finds it wherever the tests run — locally, in CI, from a packaged test JAR. Path rules follow normal classpath-resource semantics: - `"/data/users.csv"` — leading slash, absolute from the classpath root. - `"users.csv"` — no leading slash, resolved **relative to the package of the test class**, so a test in `com.acme.billing` looks for `com/acme/billing/users.csv`. Most teams use absolute paths to avoid the surprise in the second rule. **`files`** goes through the filesystem. A relative path is resolved against the JVM's working directory, which is not the same in an IDE run, a Gradle run and a Maven run. Use it only for data that is deliberately outside the build — a generated fixture, a large corpus checked out separately, a path supplied by the environment. Otherwise prefer `resources`. Both attributes are arrays, so you can list several sources; their rows are concatenated in order, and each file has its own `numLinesToSkip` applied. If a named resource or file does not exist, the test fails with a clear "could not find" style error rather than silently running zero invocations. ## numLinesToSkip Real CSV files usually start with a header row: `income,region,expectedTax`. That line is not data — attempting to convert `income` into an `int` would fail. `numLinesToSkip` tells JUnit how many leading lines to discard, and the default is `0`, meaning *no* skipping. So the near-universal setting for a human-readable file is `numLinesToSkip = 1`. It is a count of lines, not a header parser: `numLinesToSkip = 2` drops two lines, whatever they contain. It applies per file when several are listed. If you forget it, the symptom is a first invocation that fails with an argument-conversion error quoting your column names — an easy diagnosis once you have seen it. ## The other attributes - **`quoteCharacter`** — defaults to the **double quote** `"` for `@CsvFileSource`. This differs from inline `@CsvSource`, whose default is the single quote `'`. Files therefore follow ordinary CSV convention; a value containing a comma is written `"United States of America"`. Doubling the quote (`""`) escapes a literal one. - **`delimiter` / `delimiterString`** — a single char or a multi-char separator; set one, not both. Useful for tab- or pipe-separated files. - **`nullValues`** — tokens such as `N/A` or `NIL` converted to `null`. Very valuable here, because files are often authored in a spreadsheet where "nothing" gets typed several different ways. The unquoted-empty-means-null and quoted-empty-means-`emptyValue` rules of `@CsvSource` apply identically. - **`encoding`** — defaults to `"UTF-8"`. Set it when the file was exported from a tool that writes something else; a mis-set encoding shows up as mangled non-ASCII values rather than an error. - **`lineSeparator`** — defaults to `"\n"`. Files with CRLF endings normally still work because the carriage return is trimmed as whitespace, but if you are matching exact string values, be aware of it. - **`maxCharsPerColumn`** — a parser guard with a modest default (4096 characters). A single very long cell — a big JSON payload, a long SQL statement — trips it with an explicit error telling you the value is too long; raise the attribute, or accept the hint that the fixture belongs in `@MethodSource` reading its own file. - **`ignoreLeadingAndTrailingWhitespace`** — trimming, on by default, same as inline. ## When to prefer it Move the table into a file when it grows past comfortable reading in an annotation, when non-developers (QA, domain experts, product) should be able to edit the cases, when the same data feeds more than one test, or when the file is itself an artefact — a regulator-supplied bracket table, an exported golden set. Keep it inline when the cases are few and their meaning is local to the test. The cost of externalising is that a reviewer now reads the test and the data in two places, and the file must stay in the classpath. That is a fair trade for a hundred rows and a bad one for three.
- Where should the CSV file live in a typical Gradle or Maven project, and why does that matter?Under `src/test/resources`, so the build copies it onto the test classpath and `resources = "/name.csv"` finds it everywhere — IDE, local build and CI. Using `files` with a relative path instead binds the test to the working directory, which differs between an IDE run and a build-tool run, producing the classic passes-locally-fails-in-CI report.
- What is the symptom of forgetting numLinesToSkip on a file that has a header row?The first invocation tries to convert the header text into the declared parameter types and fails with an argument-conversion error quoting your column names, for example "income" cannot be converted to int. Setting `numLinesToSkip = 1` removes it. The attribute defaults to 0, so headers are never skipped automatically.
saying these in an interview costs you the question
- Believing a header row is detected and skipped automatically
- Mixing up the quote characters — assuming @CsvFileSource defaults to the single quote
- Using files with a relative path for data that ships with the project
- Expecting a non-existent resource to yield zero invocations instead of a failure
- Thinking each listed resource needs its own annotation rather than being an array element