skip to content

A Kotest assertion using shouldContainExactly on a repository result passes locally and fails intermittently in CI. How do you decide between tightening, weakening or restructuring the assertion?

level: seniorimportance: should knowfreq 26%

answer

  1. Contract first, matcher second
  2. Ordered contract -> flake is a real bug
  3. InAnyOrder still pins size + duplicates
  4. Split: contents any-order + shouldBeSortedWith
  5. Never size + contains

basics

~20 s

Decide from the contract, not from the flake. If ordering is guaranteed (an explicit ORDER BY), the flake is a real bug — keep shouldContainExactly and fix the source. If ordering is genuinely undefined, switch to shouldContainExactlyInAnyOrder, which still pins size and duplicates. Never fall back to size-plus-contains.

solid answer

~60 s

The intermittent failure is information: it says the source's ordering is not what the test assumed. Step one is to establish the contract. Does the query have an explicit deterministic ordering, or is the order an accident of the storage engine, a `Set`, a `HashMap` iteration, or parallel work? If the contract promises order, then `shouldContainExactly` is correct and the defect is in the production code — weakening the assertion buries a real bug that users will hit as inconsistent pagination. If the order genuinely is undefined, express that: `shouldContainExactlyInAnyOrder`. It is the honest assertion and still pins size and duplicate counts, so it is only one notch weaker. What I would not do: sort the actual result inside the test just to keep `shouldContainExactly`, because that hides the decision; nor replace it with `shouldHaveSize` plus a `shouldContain`, which passes for many wrong results. If ordering matters but only on a subset — ordered by score, ties arbitrary — assert the ordering with `shouldBeSortedWith` and the contents with `shouldContainExactlyInAnyOrder`, so each property is pinned separately.

code

kotlin · 9 lines
kotlin
// order is contractual: keep the strong assertion, fix the query's tie-breaker
results shouldContainExactly listOf(rowA, rowB, rowC)

// order is not contractual: state that honestly
results shouldContainExactlyInAnyOrder listOf(rowA, rowB, rowC)

// order guaranteed only by score, ties arbitrary
results shouldContainExactlyInAnyOrder expectedRows
results.shouldBeSortedWith(compareByDescending { it.score })

go deeper

for a junior

Know the two matchers and that the choice depends on whether the result has a defined order.

for a middle

Justify the choice from the contract and note that in-any-order still enforces size and duplicates.

for a senior

Lead with 'the flake is information': diagnose the ordering contract, split contents from ordering assertions where the order is partial, and reject intent-hiding workarounds.

for a principal

Make it a policy: ordering guarantees belong in the API contract and in one place in the test conventions, so teams stop rediscovering them test by test and stop trading correctness for green builds.

## The decision, not the fix This question is about judgement. Any candidate can make a flaky assertion green; the interviewer wants to hear you route through the contract first. ## Step 1: what does the code under test promise? Ask what defines the order of the result: - an explicit, total ordering in the query or in the code (`ORDER BY created_at, id`) — order **is** contractual; - an ordering with ties and no tie-breaker (`ORDER BY score` where scores collide) — order is *partially* contractual; - no ordering clause at all, a `Set`, a `HashMap` traversal, results assembled from concurrent work — order is **not** contractual and any stability you observed locally was luck. Only the third case justifies weakening. In the first, an intermittent failure is a genuine defect: the code promises an order it does not deliver, and users would see rows shuffle between page loads. Weakening the test there converts a caught bug into a shipped one. ## Step 2: assert exactly what is contractual Map each case onto a matcher. **Fully ordered contract.** ```kotlin results shouldContainExactly listOf(rowA, rowB, rowC) ``` Keep it. Fix the query — usually by adding a deterministic tie-breaker such as the primary key to the ordering. **No ordering contract.** ```kotlin results shouldContainExactlyInAnyOrder listOf(rowA, rowB, rowC) ``` This stays strong on the axes that still matter: same size, same elements, same duplicate counts. It is a one-notch weakening, not a surrender. **Partial ordering — ordered by a key with ties.** ```kotlin results shouldContainExactlyInAnyOrder expectedRows // contents results.shouldBeSortedWith(compareByDescending { it.score }) // the guaranteed order ``` Two assertions, each pinning one real property. This is the answer that distinguishes a senior candidate: it neither over-specifies (demanding an arrangement the system never promised) nor under-specifies. ## Step 3: the things not to do **Sorting the actual value to force a match.** ```kotlin results.sortedBy { it.id } shouldContainExactly expected.sortedBy { it.id } ``` It works, but the test now silently means "any order" while looking like an ordered assertion — the reader has to notice the `sortedBy` to understand the contract. `shouldContainExactlyInAnyOrder` states it in the matcher name. **Falling back to a weak pair.** ```kotlin results shouldHaveSize 3 results shouldContain rowB ``` This passes for any three-element result containing `rowB` — wrong elements, duplicates, missing rows all go undetected. It is the most common way a flaky test becomes a useless test. **Retrying or ignoring.** Marking the test flaky-tolerant hides an ordering defect and, worse, trains the team to ignore that class of failure. ## Step 4: consider the fixture, not just the matcher Sometimes the right fix is upstream of the assertion entirely: - If nondeterminism comes from shared state between tests (a database not reset, a static cache), the order flake is a symptom of test isolation, and no matcher change fixes it. - If the result is genuinely unordered but the *test* would be clearer over a canonical form, map to a projection first: `results.map { it.id } shouldContainExactlyInAnyOrder listOf(1, 2, 3)`. Comparing small projections gives sharper failure output than comparing big entities. ## What good looks like in the answer 1. Treat the flake as a question about the contract, not an inconvenience. 2. Name `shouldContainExactly` versus `shouldContainExactlyInAnyOrder` precisely, including that the latter still enforces size and duplicate counts. 3. Offer the split assertion — contents in any order plus `shouldBeSortedWith` for the guaranteed ordering key — for the partial-order case. 4. Explicitly reject the size-plus-contains fallback and in-test sorting as intent-hiding. 5. Mention that the flake may not be about ordering at all, but about test isolation.

  • Why is shouldHaveSize plus shouldContain a bad replacement for a flaky shouldContainExactly?
    Because it is dramatically weaker: any collection of the right length containing the one named element passes, so wrong elements, missing rows and duplicates all escape. `shouldContainExactlyInAnyOrder` gives up only the ordering axis while still enforcing size, membership and duplicate counts, which is almost always the correct amount of weakening.
  • What is wrong with sorting the actual collection in the test to keep shouldContainExactly?
    It produces a test whose stated matcher claims an ordering guarantee while the `sortedBy` call quietly removes it, so a reader has to reconstruct the real contract from the plumbing. `shouldContainExactlyInAnyOrder` puts the same meaning in the matcher name, which is self-documenting and gives a failure message aligned with what is actually asserted.
  • The flake persists after switching to shouldContainExactlyInAnyOrder. What next?
    Suspect test isolation rather than ordering: leftover rows from a previous test, a shared cache, or a container reused across specs will change the size and membership, which the in-any-order matcher still enforces. Investigate the fixture lifecycle and shared state before touching the assertion again.

saying these in an interview costs you the question

  • Weakening the assertion without first asking whether ordering is contractual
  • Replacing an exact assertion with size plus a single contains
  • Sorting the actual result in the test and leaving the matcher name implying order
  • Marking the test as tolerated-flaky instead of diagnosing it
  • Assuming any intermittent collection failure must be about ordering rather than shared state

context