skip to content

How would you design a shared set of domain-specific Kotest matchers for a team — the naming and negation surface, how the assertions are exposed to test authors, and what keeps failure output good as the set grows?

level: principalimportance: should knowfreq 18%

answer

  1. be… / have… predicate phrases, parameterise over proliferate
  2. always ship shouldX and shouldNotX, return the receiver
  3. delegate to should/shouldNot — never throw inside a matcher
  4. contramap for reuse, neverNullMatcher for nullable carriers
  5. test messages in both polarities; wording is production copy

basics

~20 s

One function per rule returning a Matcher<T>, named as a predicate phrase (beShippable()). Expose paired shouldX / shouldNotX extensions that delegate to should / shouldNot and return the receiver. Never throw from inside a matcher, keep matchers pure, and review failure wording like production copy.

solid answer

~50 s

I'd keep three layers. **Rules**: one function per domain rule returning `Matcher<T>`, named as a predicate phrase — `beShippable()`, `haveStatus(PAID)` — with both messages worded honestly and the actual value rendered. **Surface**: for each rule, a pair of extensions `Order.shouldBeShippable()` / `shouldNotBeShippable()` that delegate to `should` / `shouldNot` and return the receiver so calls chain. Never throw an `AssertionError` inside the matcher — routing through the entry points is what preserves contextual clues and lets failures be collected rather than thrown inside soft-assertion blocks. **Reuse**: adapt with `contramap` rather than duplicating predicates, wrap nullable carriers with `neverNullMatcher`, and prefer a hand-written composite over deep `and`/`or` chains when the rule has a domain name, because composition never merges messages. Operationally: put them in a shared test-fixtures module, unit-test each matcher in both polarities, and review message wording as seriously as production log lines — that text is the whole diagnostic for everyone who did not write the test.

code

kotlin · 19 lines
kotlin
// layer 1 - the rule
fun beShippable(): Matcher<Order> = Matcher { order ->
    val problems = buildList {
        if (order.items.isEmpty()) add("has no items")
        if (order.address == null) add("has no shipping address")
        if (!order.paid) add("is unpaid")
    }
    MatcherResult(
        problems.isEmpty(),
        { "order ${order.id} should be shippable but ${problems.joinToString(" and ")}" },
        { "order ${order.id} should not be shippable" },
    )
}

// layer 2 - the surface, both polarities, receiver returned for chaining
fun Order.shouldBeShippable(): Order { this should beShippable(); return this }
fun Order.shouldNotBeShippable(): Order { this shouldNot beShippable(); return this }

infix fun Order.shouldHaveStatus(expected: Status): Order { this should haveStatus(expected); return this }

go deeper

for a junior

Focus on the mechanics: a factory returning Matcher<T>, plus a shouldX / shouldNotX pair delegating to should and shouldNot.

for a middle

Add naming conventions, returning the receiver, and why matchers must be pure and word both polarities.

for a senior

Bring in reuse via contramap and neverNullMatcher, hand-written composites for message quality, and unit-testing the matchers' own messages.

for a principal

Argue the scope trade-off explicitly — a small owned vocabulary versus an unlearnable internal DSL — and set the team rules: where matchers live, who reviews wording, and when a matcher is not warranted.

## Why a matcher library is a design problem, not a utility drawer A custom matcher is read by two audiences: the test author choosing it from autocomplete, and the engineer three months later staring at a red CI job with no source open. The first cares about naming and the shape of the call; the second cares only about the message. Most homegrown matcher sets serve the first and neglect the second. ## Layer 1 — rules as matcher factories Write one function per rule that returns a `Matcher<T>`: - **Name it as a predicate phrase** so it reads inside an assertion: `beShippable()`, `haveStatus(PAID)`, `containOnlyDigits()`. Kotest's own convention is `be…` for a property of the value and `have…` for something it contains. - **Parameterise instead of proliferating.** `haveStatus(status)` beats `beShipped()` + `bePaid()` + `beCancelled()` — one rule, one message template, one test. - **Keep matchers pure.** No I/O, no mutation, no consuming iterators. Composition and inversion may call `test()` more than once on the same value. - **Word both polarities honestly.** The negated message is the entire output of every `shouldNot` failure; a copy of the positive text is a latent bug that surfaces the first time somebody asserts a negative. - **Render the actual value.** "order 4711 should be shippable but is unpaid" beats "not shippable". ## Layer 2 — the assertion surface Test authors should not be writing `should beShippable()` everywhere unless you want them to. Offer paired extensions: ```kotlin fun Order.shouldBeShippable(): Order { this should beShippable(); return this } fun Order.shouldNotBeShippable(): Order { this shouldNot beShippable(); return this } ``` Design points: - **Always ship both directions.** A one-sided surface pushes people back to raw `shouldNot`, defeating discoverability. - **Return the receiver.** It mirrors Kotest's own assertions and lets calls chain. - **Delegate to `should` / `shouldNot`; never throw directly.** Those entry points are where contextual clues are attached and where a failure is *collected* rather than thrown when the assertion runs inside a soft-assertion block. A matcher or extension that throws its own `AssertionError` silently opts out of both, and cannot be negated or composed. - **Use `infix` only for one-argument forms** (`order shouldHaveStatus PAID`); zero-argument extensions cannot be infix. - **Keep the matcher factory public too.** Assertions are for test bodies; the raw matcher is what inspectors, collection matchers and `and`/`or` need. ## Layer 3 — reuse without duplication - **`contramap`** adapts an existing matcher onto a new carrier type via a projection, so a string-level rule serves every domain type that holds such a string. Budget for re-wording: the inherited message describes the projected value and loses the owner's identity. - **`neverNullMatcher`** lifts the null check out of the rule, turning a null into a reported failure instead of an NPE thrown from inside `test()`. Decide once, as a team, that null violates a content rule rather than vacuously satisfying it. - **`invert()`** when a negation must be a value — an operand of `and`/`or`, or something handed to an API taking a `Matcher<T>`. - **Prefer a hand-written composite** to a deep `and`/`or` chain for any rule with a domain name. Composition propagates one sub-result verbatim and short-circuits, so it can never say "unpaid **and** missing an address"; a hand-written matcher collects every violation and names them all. ## Keeping quality as the set grows - **Unit-test the matchers themselves**, both polarities, asserting on the message text. This is the only thing that catches a copy-pasted negated message. - **Place them in a shared test-fixtures source set or module** so both application and integration tests use one copy — and so the matchers never leak into production code. - **Review wording like production copy.** Agree on a template: `<subject> should <rule> but <observed>`. Consistency matters more than eloquence. - **Cap the surface.** A matcher earns its place when the rule appears in three or more tests or when the naive assertion produces an unreadable failure. Otherwise a plain expression is clearer than a bespoke vocabulary nobody remembers. - **Watch for matchers that encode business logic.** If a matcher reimplements the rule the production code implements, the test can pass while both are wrong together. Matchers should observe state, not recompute policy. ## The trade-off to state out loud A rich matcher vocabulary makes tests read like specifications and makes failures self-explanatory — at the cost of an internal DSL every newcomer must learn, and of drift if nobody owns it. The judgement call is scope: a small, well-named, well-tested set around the core aggregates pays for itself; a matcher per predicate in the codebase does not.

  • Why insist that the extension delegates to should/shouldNot rather than throwing AssertionError itself?
    Those entry points are the framework's failure pipeline: they attach any contextual clues in scope and decide whether the failure is thrown immediately or collected, which is what makes assertions inside a soft-assertion block report together. An extension that throws on its own bypasses the pipeline, so it loses context and short-circuits blocks that were meant to gather every failure. It also cannot be negated or composed, because there is no MatcherResult to work with.
  • When is a custom matcher the wrong call?
    When the rule appears once or twice, when a plain expression already fails readably, or when the matcher would reimplement the production logic it is meant to check — in that last case a bug in the logic is mirrored in the matcher and the test passes anyway. A matcher earns its place through repetition plus a measurable improvement in failure output, not through elegance.

saying these in an interview costs you the question

  • Shipping only the positive assertion and leaving the negative direction to raw shouldNot
  • Throwing AssertionError from inside the matcher or the extension
  • Building named domain rules out of deep and/or chains and expecting merged messages
  • Treating failure wording as an afterthought instead of the library's main product
  • Letting matchers recompute business logic so a bug is mirrored on both sides of the test

context