skip to content

What is a characterization test, how do you write one for code you do not understand, and what should you do when it reveals behaviour that looks like a bug?

level: middleimportance: must knowfreq 58%

answer

  1. Records what code does, not what it should do
  2. Assert something wrong, read failure, paste real value
  3. Coverage + boundaries pick the inputs
  4. Bugs get characterized, not silently fixed
  5. Golden master needs determinism

basics

~20 s

A characterization test pins down what code actually does today, not what it should do. You call the code, assert a deliberately wrong value, read the failure message to learn the real value, and change the assertion to match. If the real behaviour looks like a bug, you still record it, then ask before changing it.

solid answer

~50 s

A characterization test (also called an approval or golden-master test) documents existing behaviour so you can detect unintended change while refactoring. Technique: write a test that instantiates the code, calls it, and asserts something you know is false; run it; the failure message reveals the actual value; paste that value into the assertion. Repeat, using coverage and input-space reasoning to add cases that exercise each branch. The tests need not be pretty or express intent — their only job is change detection. Crucially they are not specifications: they lock in bugs as faithfully as correct behaviour. When you find behaviour that looks wrong, do not silently fix it — someone may depend on it. Record it, mark it, and escalate to whoever owns the product decision; fixing becomes a separate, deliberate change with its own test. For wide output surfaces, a golden-master approach (serialize the whole output, diff against an approved snapshot) covers more behaviour per unit of effort.

code

pseudocode · 12 lines
pseudocode
// Step 1: assert something you know is false
test "characterize invoiceTotal" {
  result = calculator.invoiceTotal(order)
  assertEquals(-999, result)      // deliberately wrong
}
// Run -> failure says: expected -999 but was 4210

// Step 2: record reality (even if 4210 looks off by a cent)
test "characterize invoiceTotal" {
  result = calculator.invoiceTotal(order)
  assertEquals(4210, result)      // TODO: rounding looks wrong; asked product, ticket ABC-12
}

go deeper

for a junior

Define it as a test that records current behaviour, and describe the assert-wrong-then-paste-the-real-value loop.

for a middle

Add input selection via coverage and boundaries, sensing of side effects via fakes, and the rule that suspected bugs get characterized and escalated rather than fixed inline.

for a senior

Discuss golden-master/approval testing for wide outputs, determinism requirements, assertion granularity trade-offs, and migrating characterization tests into intention-revealing specification tests over time.

for a principal

Frame it as risk management and change-failure-rate reduction: where to place the characterization boundary, Hyrum's Law implications for quirks, and the policy that behaviour changes ship separately from refactorings so the audit trail stays readable.

## What it is A **characterization test** is a test written to *characterize* (describe) the behaviour a piece of code has right now. It is the tool that resolves the practical problem of "I must change this code but I do not know what it is supposed to do". Contrast with the two other test intents: - A **specification test** asserts what code *should* do, derived from requirements. It can legitimately fail against buggy code. - A **regression test** guards a previously fixed defect. - A **characterization test** asserts what code *does* do. By construction it passes on day one. If it fails, either you changed behaviour or the environment did. Alternative names for the same idea: **approval test**, **golden master test**, **snapshot test**, **pinning test**. ## The mechanical recipe Feathers describes an almost thoughtless loop, which is the point — it works even when you cannot read the code: 1. **Get the code into a test harness.** Instantiate the class or call the function from a test. This is usually the hard part and may require dependency-breaking first (see seams, sprout, wrap). 2. **Write a test that asserts something you know is wrong.** For example, assert the result equals `"XYZZY"` or `-1`. 3. **Run it and read the failure.** The assertion message tells you the real value: `expected "XYZZY" but was 4210`. 4. **Change the assertion to the observed value.** The test now passes and encodes real behaviour. 5. **Repeat for more inputs** until you have covered the behaviour you are about to disturb. ## Choosing inputs Random inputs are weak. Pick inputs deliberately: - **Branch coverage as a guide.** Run coverage; every uncovered branch in the region you will touch is a behaviour you cannot detect breaking. Craft an input that reaches it. - **Boundaries.** Empty collection, single element, maximum size, zero, negative, null/absent, just-below and just-above every literal threshold you can see in the code. - **Error paths.** Assert the exception type and, if it is part of the contract, the message. - **Side effects, not just returns.** If the method mutates state, writes to a collaborator, or emits an event, capture that too — this is *sensing*, usually done with a fake or spy collaborator. ## The golden-master variant When the output is large (a rendered report, a generated file, an object graph), asserting field by field is slow. Instead: 1. Serialize the entire output to text (JSON, CSV, formatted dump) in a stable, deterministic order. 2. Store an **approved** copy on disk on the first run, after a human eyeballs it. 3. Later runs diff actual against approved; any difference fails the test and is shown as a text diff. This gives enormous behavioural coverage per test, and works well for legacy batch jobs and report generators. Cost: the diff tells you *that* something changed, not *why*, and approvals can be rubber-stamped without reading. ### Determinism is mandatory Golden masters break on anything non-deterministic: timestamps, random values, UUIDs, hash-ordered maps, absolute file paths, locale-sensitive number/date formatting, thread interleaving. You must inject a fixed clock/seed or scrub those fields before comparison, otherwise the suite becomes flaky and people start ignoring it. ## What to do when the behaviour looks wrong This is the most-tested judgement point in interviews. The rule: **the characterization test records reality, including bugs.** Reasons: - Callers may already depend on the quirk (Hyrum's Law: with enough users, every observable behaviour of your system will be depended on by somebody). - You cannot tell, from inside a refactoring, whether the quirk is load-bearing. - Mixing a behaviour change into a refactoring destroys the refactoring's safety property: when the suite goes red you no longer know whether you broke something or intentionally changed it. So: write the assertion capturing the odd value, add a comment or a distinctive test name marking it as suspected-incorrect, raise a ticket/ask the product owner, and if the fix is approved, do it as a **separate commit** where the changed characterization test is the visible diff. ## Trade-offs and limitations - **They are not documentation of intent.** A future reader sees `assertEquals(4210, result)` with no explanation of why 4210 is right. Over time, replace valuable ones with intention-revealing specification tests as you learn the domain. - **They can lock in the wrong thing.** If you characterize at too coarse a boundary (e.g. through the HTTP layer only), a refactoring can preserve the outer output while corrupting an internal invariant. - **They cost real time** on genuinely hostile code, but the cost is bounded to the change point rather than the whole system. - **Over-specification.** Asserting a full toString dump makes every cosmetic change fail. Match assertion granularity to what you intend to keep stable. ## Where they sit in the workflow In Feathers' Legacy Code Change Algorithm the order is: identify change points, find test points, break dependencies, **write characterization tests**, then make changes and refactor. They are the safety net that makes the last step non-scary.

  • How is a characterization test different from a unit test written with TDD?
    Intent and failure semantics. A TDD test states required behaviour and is written red, before the code exists; it is a specification. A characterization test is written against existing code, is expected to pass immediately, and states only 'this is what happens today'. A TDD test failing means the code is wrong; a characterization test failing means the behaviour moved.
  • You cannot even construct the class to characterize it — the constructor opens a database connection. What now?
    Break the dependency first with the smallest safe move: extract the connection creation into an overridable factory method and subclass it in the test, extract an interface and pass a fake in, or use a link/build-time seam to substitute the driver. Alternatively use Sprout Method so the new logic lives in a testable unit even if the host class stays untestable for now.
  • How many characterization tests are enough?
    Enough to cover the behaviour your planned edit could disturb — typically driven by branch coverage of the change region plus its callers' observable effects. Full-system characterization is not the goal.

It is like taking a plaster cast of a machine part before you start filing it: the cast is not a blueprint of the ideal part, it is a record of the exact shape you had, so afterwards you can tell whether you removed anything you did not mean to.

saying these in an interview costs you the question

  • Fixing an obviously-wrong value while writing the characterization test, in the same commit as a refactoring
  • Calling characterization tests 'specifications' or treating them as requirement documentation
  • Using random or arbitrary inputs instead of boundary- and branch-driven inputs
  • Golden masters containing timestamps, UUIDs or unordered map output, then blaming flakiness on the tool
  • Asserting only 'no exception thrown', which detects almost no behaviour change
  • Deleting a characterization test because it 'asserts a bug'

context