A new unit test passes on its first run, before the code it drives exists. How do you diagnose that?
answer
- Falsified prediction, not good luck
- Force it red before anything else
- Executed at all, or just green?
- Where did the expected value come from?
- Sometimes the behavior already exists
basics
~20 sTreat an unexpected green as a broken test. Force it red by changing the expected value to something certainly wrong; if it still passes, the case is not executed, the assertion is unreachable, or the comparison is vacuous.
solid answer
~50 sFirst establish whether the case runs at all: change the expected value to something that cannot be right and re-run. If it stays green, the runner never executed it — a discovery rule, a skip marker, a tag filter, or a filtered command. If it now goes red, the case runs and the question moves to the assertion: is the expected value computed by the same code the test exercises, does the fixture write back the value the assertion reads, does a stand-in answer the check directly, or is a numeric tolerance wide enough to swallow the difference? A third possibility deserves attention rather than embarrassment: the behavior may already exist somewhere, in which case you have discovered your task is partly done and should decide whether to keep the case as a characterization test.
code
pseudocode · 6 linestest "income with a grouping separator is rejected":
result = validateIncome("1 247,35")
assert result.message == "IMPOSSIBLE_SENTINEL_9137"
# still green -> the case is never executed
# now red -> the case runs; inspect the assertiongo deeper
Remember the reflex: an unexpected pass means something is wrong with the test, not that you got lucky. Recall the one-edit check — change the expected value to something impossible and run again.
Be ready to walk the diagnosis in order and explain what each outcome rules out: execution first, then reachability, then whether the expected value was derived from the very code being exercised.
Demonstrate that you look for the same failure mode across a whole suite, not one case: derived oracles, doubles that answer assertions directly, and wide numeric tolerances are systemic patterns that need review rules, not one-off fixes.
Own the detection strategy at scale. Argue for signals that surface never-failing cases across many suites, for review guidance on where expected values may come from, and for how much analysis effort a codebase should spend on tests that may be guarding nothing.
## An unexpected green is a defect report about your test In the red step you have a prediction: this case should fail, because the behavior it describes does not exist. A green run falsifies the prediction. Something you believe about the system, the test, or the runner is wrong, and it is much cheaper to find out now than after the case has been sitting in the suite for a year pretending to guard something. Work the diagnosis cheapest-first, because the cheap checks eliminate whole categories. ## Step 1 — is the case executed at all? Change the expected value to something that cannot possibly be produced, and run again. - **Still green.** The case never ran. Look for: a file or case name that does not match the runner's discovery rules; a skip or ignore marker; a tag or category filter in the command actually being used; a case defined inside a scope that is never collected; a build step that did not recompile what you edited so an older version of the test executed. - **Now red.** The case runs. Restore the real expectation and continue to step 2. This single check is the highest-value move in the whole diagnosis, and it takes one edit and one run. It also tells you the sharpest fact available: whether you are debugging the runner or the assertion. ## Step 2 — is the assertion reachable and meaningful? If the case runs but still passes, work down this list: - **Unreachable assertion.** The body returns early, takes a branch the fixture does not satisfy, or wraps the action in a construct that swallows the failure. An assertion after a loop that never iterates is the classic version. - **Derived oracle.** The expected value is produced by calling the same function, formatter or mapper the production path calls. Both sides move together, so the comparison can never disagree. This is the single most common source of a silently useless test. - **The fixture supplies the answer.** Setup writes a record and the assertion reads the same record back through a layer that does no transformation, so the test measures the fixture. - **A stand-in answers the check.** A configured double returns exactly the value asserted on and the real collaborator is never reached, so the case passes whatever the real one would do. - **Coincidental default.** The system returns an empty collection, a zero or an absent value, and the expectation happens to be the same. The case is passing by accident and will keep passing when the behavior is wrong in a different way. - **Tolerance too wide.** A numeric comparison with a margin larger than the error being tested cannot see the error. - **Asynchronous timing.** The assertion evaluates before the action completes, so it inspects an initial state that happens to satisfy it. ## Step 3 — maybe the behavior really is there Sometimes the test is sound and the system already does the thing. That is a genuine and useful discovery: the feature exists elsewhere, a default already covers the case, or another team shipped it. Then decide deliberately — keep the case as a characterization test that pins existing behavior, or delete it as a duplicate of a case that already covers the path — and revisit your task, because part of the work you planned may be unnecessary. ## A worked example On a tax-filing wizard, a developer writes a case for a new step that should reject a declared income entered with a locale-dependent thousands separator. The case passes on the first run. Forcing the expected value to a nonsense string still leaves it green, which points at execution rather than the assertion; the file sits one directory outside the path the runner scans, so it is compiled but never collected. Moved into the scanned path, the case runs — and now passes for a second reason: the expected rejection message is built with the same localisation lookup the validator uses, so the comparison agrees under every locale. Two independent defects in one case, both invisible without the forced red, and both found in about three minutes. ## The habit this leaves you with The general principle is that you never trust a green you have not earned. Any time a test's meaning is in doubt — after editing its fixtures, after a refactor of its doubles, after inheriting it — the way to restore confidence is the same: make the production code wrong on purpose, watch the case go red, then put the code back. Everything else is inference.
- You force the expectation to a nonsense value and the case still passes. What is your next move?Stop reasoning about the assertion and prove execution. Add an unconditional failure at the top of the case body and re-run; if the run is still green the case is not being collected. Then check discovery rules, skip markers, tag filters, the exact command being run, and whether the build recompiled your edit.
- How do you spot an oracle that is derived from the system under test during review?Read the expected side of every assertion and ask where its value came from. If it is produced by calling production code — a formatter, a mapper, a serializer, a constant owned by the production module — the comparison can never disagree with the system. Expected values should be literals, or built by logic the test owns.
- The test turns out to be correct and the behavior already exists. Is the case wasted work?No, but it needs a decision. Keep it as a characterization test if it pins behavior nothing else covers, and delete it if an existing case already exercises the same path. Either way, revisit the task: discovering the feature is already present is more valuable than the case itself.
saying these in an interview costs you the question
- Celebrates the green and moves on
- Adds more assertions instead of forcing a red
- Assumes the runner always executes every case written
- Never questions where the expected value came from
- Widens a tolerance until the case behaves
- Deletes the case rather than diagnosing why it passed