skip to content

A Kotlin service suite leans heavily on MockK spies — spyk on production classes with one or two members stubbed, plus recordPrivateCalls used to verify internal helpers. As the person setting the policy, how would you decide what stays, and which MockK mechanics drive your reasoning?

level: principalimportance: nice to knowfreq 18%

answer

  1. grade by mechanic, not taste
  2. copy mild, constructor-free obscure, strings worst
  3. stub public members only
  4. private-call assertion → refactor ticket
  5. spies bridge legacy, injection for new code

basics

~20 s

Judge each spy by the mechanic it depends on: state copying, call-through executing real code, and string-named private-call verification. Keep spies at stable public seams; retire private-call assertions and constructor-free spies as refactors make injection possible.

solid answer

~50 s

I would grade spies by which MockK mechanic each one leans on. **Acceptable**: a spy on a fully constructed object where one *public* member is stubbed and assertions are on outcomes. The mechanics involved — state copy at creation, call-through for the rest — are stable, and the test still exercises real code. **Watch**: `spyk<T>()` with no constructor run. Fields are at defaults, so the test only holds together as long as nothing real touches state; it fails as an obscure NPE when someone adds a field read. **Retire**: `recordPrivateCalls` plus `invoke "name" withArguments` assertions. Member names are strings, so renames and signature changes break tests at runtime with messages that describe MockK, not intent — a tax on every refactor. Policy: spies as a bridge for code you cannot yet inject into; each private-call assertion carries a named owner and a refactor ticket; new code gets injected collaborators instead.

go deeper

for a junior

Say the simple version: prefer injecting collaborators and mocking them; a spy is for code you cannot change yet.

for a middle

Name the concrete mechanics — state copy, call-through, string-named private calls — and which ones you would allow.

for a senior

Give a graded policy with reasoning per mechanic and a migration path anchored in the classes, not the tests.

for a principal

Own the trade explicitly: spies convert a design problem into a test problem, so make each one a named, ticketed decision with an exit plan and a visible trend.

## Why this needs a policy at all Spies are seductive: they let you test a class you cannot otherwise construct or isolate, without changing production code. The cost is not paid at write time — it is paid every time someone refactors. A policy is worth writing because the cost is invisible in code review and very visible six months later. ## Grade by mechanic, not by taste Each spy usage depends on specific MockK mechanics, and each mechanic has a different failure profile. **1. State copying (`spyk(obj)`).** MockK copies the object's field values into its own instrumented instance; the copy is shallow and taken at creation time. Failure profile: mild and local. Tests break when someone wires state after spy creation, and the fix is obvious once known. Low tax — this usage can stay. **2. Constructor-free creation (`spyk<T>()`).** No constructor, initialiser or `init` block runs, so fields hold `null`/`0`/`false`. Failure profile: delayed and obscure. The test passes until someone adds a field read to a method the test reaches, then it throws an NPE or `lateinit property ... has not been initialized` from inside production code. Medium tax — allowed only for genuinely stateless classes, and worth a comment saying why. **3. Call-through.** Anything unstubbed runs real code, and a stub that does not match its arguments silently falls through to that real code. Failure profile: tests that quietly do more than they claim — real computation, occasionally real side effects. Medium tax — control it by keeping the stubbed surface small and explicit, and by reading a growing stub list as a signal to switch to a mock plus a narrower unit under test. **4. Private-call recording (`recordPrivateCalls = true`, `invoke "name" withArguments listOf(...)`).** Members are addressed by string. Failure profile: worst of the set. A rename compiles cleanly and fails at runtime; adding a parameter to a private helper — normally a free refactor — breaks tests; and the assertion states "the implementation does it this way", which is not a behavioural claim at all. High tax. ## The policy I would actually write 1. **Default is injection, not spying.** New code takes its collaborators as constructor parameters, so tests use plain mocks and no spy is needed. 2. **Spies are a bridge, not a destination.** A spy is acceptable where the class cannot yet be injected into. Each one is a marker on the map of code that has not been refactored. 3. **Stub public members only.** If the thing you need to replace is private, that is the design telling you it wants to be a collaborator. 4. **Private-call verification is quarantined.** Every `invoke "name"` assertion carries a linked refactor ticket and an owner; when the helper becomes an injected collaborator, the dynamic call is deleted with it. Trend toward zero. 5. **Prefer stubbing a private call to verifying one.** Stubbing buys isolation and breaks only when the seam moves; verifying pins the implementation shape and breaks on any refactor. 6. **No constructor-free spies on stateful classes.** If constructing the class in a test is painful, fix the construction rather than skipping it. 7. **Assert outcomes over interactions** wherever the outcome is observable; spy verification is a fallback for effects with no observable result. ## How I would migrate Inventory first: count `spyk(` and `recordPrivateCalls` occurrences and see how they cluster. Clusters usually map onto two or three god-classes. Attack those classes rather than the tests — extracting one collaborator often deletes a dozen dynamic calls at once. Add the policy to review guidance so the count stops growing while the refactor proceeds. ## What I would tell the team The honest framing is that a spy converts a design problem into a test problem, and private-call verification converts it into a *string-typed* test problem. That is sometimes the right trade — the alternative is often no test at all — but it should be a decision with a name on it and an exit plan, not the default idiom of the suite.

  • Someone argues that verifying a private helper is fine because the test is green today. What is your counter?
    Green today is not the cost centre. The member is named by a string, so the very refactors you want to encourage — renaming a helper, adding a parameter — break the test at runtime with a message about a missing member rather than about behaviour. Over a suite of hundreds of such assertions, that becomes a standing tax on every refactor and an argument against improving the code.
  • How would you keep spy usage from creeping back after the migration?
    Make the count visible: track occurrences of spyk and recordPrivateCalls over time and review changes that add them, requiring a one-line justification. Pair that with making the good path easy — constructor-injected collaborators and mockable interfaces — so the spy is no longer the path of least resistance.

saying these in an interview costs you the question

  • Treating spies as a normal default rather than a bridge for hard-to-inject code
  • Defending private-call verification as 'stronger' testing when it only pins implementation shape
  • Ignoring that unmatched argument matchers on a spy silently execute real production code
  • Standardising on spyk<T>() to avoid awkward constructors instead of fixing construction
  • Judging spy usage by style preference rather than by the failure profile of the mechanic it relies on

context