For a large parameterized suite in Kotest, how would you decide between running cases through a `table(...)` with `forAll` and registering them with `withData` — and what convention would you set for a team?
answer
- one test (table) vs one test per case (withData)
- granularity, addressability, lifecycle, cost — all follow from that
- table: headers label failures, callbacks once, no per-row re-run
- withData: named nodes, per-case callbacks, filterable, nesting = matrix
- enforce DuplicateTestNameMode.Error; cap cross-products
basics
~20 sA table runs all rows inside one test: one report node, callbacks once, no per-row re-run, failures aggregated into one message. withData registers a test per case: individually named, reported, filtered and re-runnable. Choose by whether per-case triage matters more than compactness.
solid answer
~60 sThe decision is about **reporting granularity**, not syntax. A `table(...)` + `forAll` executes every row inside a single Kotest test. You get one report node, per-test lifecycle callbacks fire once for the whole table, no row can be re-run or filtered on its own, and per-row detail exists only inside the aggregated failure message. In exchange it is compact, strongly typed per column, and cheap — no name generation, no N nodes in the tree. `withData` registers a real test node per element. Each case is named, reported, individually re-runnable and addressable by name-based filtering, callbacks fire per case, and a nested `withData` gives you a matrix of leaves. The cost is names you must keep stable and readable, and a report tree that grows with the data. My convention: tables for small closed sets of pure input/output rows inside one behaviour; `withData` whenever cases need per-case setup, per-case triage, or will be quoted in a bug report. Cap nested cross-products, and enforce naming quality so the `withData` side stays readable.
code
kotlin · 20 lines// Table: one test node, aggregated failures, compact
test("discount tiers") {
forAll(
table(
headers("spend", "tier", "expected"),
row(0, FREE, 0),
row(10_000, PRO, 500),
),
) { spend, tier, expected -> discount(spend, tier) shouldBe expected }
}
// withData: one named node per case, per-case callbacks, individually re-runnable
data class DiscountCase(val spend: Int, val tier: Tier, val expected: Int)
context("discount tiers") {
withData(
DiscountCase(0, FREE, 0),
DiscountCase(10_000, PRO, 500),
) { (spend, tier, expected) -> discount(spend, tier) shouldBe expected }
}go deeper
Say that a table runs all rows inside one test while withData creates a test per case, and that withData names each case.
Derive the practical differences — report granularity, per-case callbacks, failure detail — from that structural difference.
Add addressability and operations: re-running and filtering a single case, triage on a red build, and the cost of large generated trees.
Commit to a default with an enforcement point (DuplicateTestNameMode.Error, matcher-in-body review rule, cross-product cap), and say honestly when the compact table wins.
## The two mechanisms, precisely **Table + `forAll`** (`io.kotest.data`): `table(headers("a","b","expected"), row(...), row(...))` executed by `forAll(table) { a, b, expected -> ... }` inside an ordinary test. All rows run; failures are collected and thrown as one aggregated error naming each failing row by header/value pairs. The whole table is **one test**. **`withData`** (`io.kotest.datatest`, artifact `kotest-framework-datatest`): registers **one test node per element**, deriving each name from the value (`WithDataTestName` → `@IsStableType` → data class `toString()` → type-derived fallback), overridable with the `nameFn` or map-of-names overloads. It may be nested to produce a cross-product of leaves. Everything else follows from "one test" versus "a test per case". ## Consequence-by-consequence comparison **Report granularity.** Table: one green or red node, and three bad rows out of twenty read as one failure. `withData`: seventeen green nodes and three red ones, which is what a triage dashboard, a flaky-test detector and a reviewer scanning a CI page actually want. **Addressability.** Table: you cannot re-run row 7 — from the IDE or from name-based filtering you can only select the enclosing test, so isolating a case means commenting rows out. `withData`: each case is a named node the IDE can re-run and filters can select, provided the names are stable. **Lifecycle.** Table: per-test callbacks fire once for the enclosing test, so any per-row setup or cleanup has to live inside the lambda by hand. `withData`: each generated case is a real test node, so per-test callbacks see it, which is what you want when a case needs a fresh fixture, a clean database row, or its own timeout. **Failure detail.** Table: the aggregated message is the only per-row information, which makes `headers(...)` and printable row values essential. `withData`: the failing node's name *is* the case identity, and the assertion message stands alone. **Typing and shape.** Table: arity-generated types give per-column static typing and compile-time arity checking, but arity is bounded and every row must share a column's type. `withData`: one value per case, so complex cases are just a data class — no arity ceiling, and the data class doubles as the name source. **Cost.** Table: no name generation, one node, negligible overhead — 200 rows is nothing. `withData`: 200 nodes with 200 computed names in the report tree, and nested matrices multiply into thousands. Usually acceptable, occasionally a real drag on IDE trees and report size. ## The convention I would set **Use a table when** the cases are a closed, small-to-medium set of pure input→expected rows for a *single* behaviour, no case needs its own setup, and nobody will need to re-run one row in isolation. Truth tables, arithmetic edge cases, parser inputs, boundary values. Requirements: always supply `headers(...)`, keep inputs first and expected last, and remember the block must assert (a boolean body silently passes). **Use `withData` when** any of these is true: a case needs per-case setup or teardown; cases will be quoted individually in bug reports or triage; the case set is long-lived and worth naming; the case is a rich object rather than a few primitives; you need a matrix across dimensions. **Enforce**: `duplicateTestNameMode = DuplicateTestNameMode.Error` in `AbstractProjectConfig` so `withData` naming defects fail the build rather than producing indexed, positional names; a review rule that every table-runner body contains a matcher call; and a cap on nested cross-product size, with the observation that a huge enumerated matrix is usually a sign the dimension should be generated rather than listed. ## Migration is cheap in one direction Going from a table to `withData` means turning each row into a data class instance — mechanical, and it improves failure output. Going the other way is equally mechanical but loses per-case reporting, so it is rarely worth doing except to collapse a noisy tree. That asymmetry is a reason to default to `withData` for anything you expect to live for years, and to keep tables for the compact, local, obviously-throwaway cases. ## The judgment being assessed A weak answer compares syntax. A strong one names the single structural difference — one test versus N tests — and then derives reporting, addressability, lifecycle and cost from it, before committing to a default and an enforcement point. It should also say honestly when the compact option wins: a twelve-row truth table as twelve report nodes is noise, not signal.
- A case in your parameterized set needs its own database fixture. Which mechanism does that force, and why?`withData`, because each generated element is a real test node, so per-test lifecycle callbacks fire for it and the isolation configuration applies per case. In a table all rows execute inside one test, so callbacks fire once for the whole table and any per-row setup has to be hand-written inside the lambda — which reintroduces exactly the leakage between cases that per-test fixtures exist to prevent.
- When is a table's single report node actually better than per-case nodes?When the cases are a compact, closed truth table for one behaviour and per-case identity carries no information a reader needs. Twelve report nodes named after boolean triples is noise on a CI page; one node called "precedence rules" with a failure message naming the bad rows is easier to scan. It is also cheaper — no name generation, no tree growth — which matters when a suite has hundreds of such tables.
saying these in an interview costs you the question
- "They're the same thing with different syntax" — the structural difference is one test versus one test per case, and everything else follows.
- "Tables give you per-row test results" — the whole table runs inside a single test node.
- "beforeTest runs per row in a table" — per-test callbacks fire once for the enclosing test.
- "Always prefer withData because more granularity is always better" — for a compact truth table, N nodes is noise and costs report size and name management.
- "You can filter or re-run an individual table row" — rows are not addressable; only the enclosing test is.