What is a characterization test, how do you write one for code you do not understand, and what should you do when it reveals behaviour that looks like a bug?
answer
- Records what code does, not what it should do
- Assert something wrong, read failure, paste real value
- Coverage + boundaries pick the inputs
- Bugs get characterized, not silently fixed
- Golden master needs determinism
basics
~20 sA characterization test pins down what code actually does today, not what it should do. You call the code, assert a deliberately wrong value, read the failure message to learn the real value, and change the assertion to match. If the real behaviour looks like a bug, you still record it, then ask before changing it.
solid answer
~50 sA characterization test (also called an approval or golden-master test) documents existing behaviour so you can detect unintended change while refactoring. Technique: write a test that instantiates the code, calls it, and asserts something you know is false; run it; the failure message reveals the actual value; paste that value into the assertion. Repeat, using coverage and input-space reasoning to add cases that exercise each branch. The tests need not be pretty or express intent — their only job is change detection. Crucially they are not specifications: they lock in bugs as faithfully as correct behaviour. When you find behaviour that looks wrong, do not silently fix it — someone may depend on it. Record it, mark it, and escalate to whoever owns the product decision; fixing becomes a separate, deliberate change with its own test. For wide output surfaces, a golden-master approach (serialize the whole output, diff against an approved snapshot) covers more behaviour per unit of effort.
code
pseudocode · 12 lines// Step 1: assert something you know is false
test "characterize invoiceTotal" {
result = calculator.invoiceTotal(order)
assertEquals(-999, result) // deliberately wrong
}
// Run -> failure says: expected -999 but was 4210
// Step 2: record reality (even if 4210 looks off by a cent)
test "characterize invoiceTotal" {
result = calculator.invoiceTotal(order)
assertEquals(4210, result) // TODO: rounding looks wrong; asked product, ticket ABC-12
}go deeper
Define it as a test that records current behaviour, and describe the assert-wrong-then-paste-the-real-value loop.
Add input selection via coverage and boundaries, sensing of side effects via fakes, and the rule that suspected bugs get characterized and escalated rather than fixed inline.
Discuss golden-master/approval testing for wide outputs, determinism requirements, assertion granularity trade-offs, and migrating characterization tests into intention-revealing specification tests over time.
Frame it as risk management and change-failure-rate reduction: where to place the characterization boundary, Hyrum's Law implications for quirks, and the policy that behaviour changes ship separately from refactorings so the audit trail stays readable.
## What it is A **characterization test** is a test written to *characterize* (describe) the behaviour a piece of code has right now. It is the tool that resolves the practical problem of "I must change this code but I do not know what it is supposed to do". Contrast with the two other test intents: - A **specification test** asserts what code *should* do, derived from requirements. It can legitimately fail against buggy code. - A **regression test** guards a previously fixed defect. - A **characterization test** asserts what code *does* do. By construction it passes on day one. If it fails, either you changed behaviour or the environment did. Alternative names for the same idea: **approval test**, **golden master test**, **snapshot test**, **pinning test**. ## The mechanical recipe Feathers describes an almost thoughtless loop, which is the point — it works even when you cannot read the code: 1. **Get the code into a test harness.** Instantiate the class or call the function from a test. This is usually the hard part and may require dependency-breaking first (see seams, sprout, wrap). 2. **Write a test that asserts something you know is wrong.** For example, assert the result equals `"XYZZY"` or `-1`. 3. **Run it and read the failure.** The assertion message tells you the real value: `expected "XYZZY" but was 4210`. 4. **Change the assertion to the observed value.** The test now passes and encodes real behaviour. 5. **Repeat for more inputs** until you have covered the behaviour you are about to disturb. ## Choosing inputs Random inputs are weak. Pick inputs deliberately: - **Branch coverage as a guide.** Run coverage; every uncovered branch in the region you will touch is a behaviour you cannot detect breaking. Craft an input that reaches it. - **Boundaries.** Empty collection, single element, maximum size, zero, negative, null/absent, just-below and just-above every literal threshold you can see in the code. - **Error paths.** Assert the exception type and, if it is part of the contract, the message. - **Side effects, not just returns.** If the method mutates state, writes to a collaborator, or emits an event, capture that too — this is *sensing*, usually done with a fake or spy collaborator. ## The golden-master variant When the output is large (a rendered report, a generated file, an object graph), asserting field by field is slow. Instead: 1. Serialize the entire output to text (JSON, CSV, formatted dump) in a stable, deterministic order. 2. Store an **approved** copy on disk on the first run, after a human eyeballs it. 3. Later runs diff actual against approved; any difference fails the test and is shown as a text diff. This gives enormous behavioural coverage per test, and works well for legacy batch jobs and report generators. Cost: the diff tells you *that* something changed, not *why*, and approvals can be rubber-stamped without reading. ### Determinism is mandatory Golden masters break on anything non-deterministic: timestamps, random values, UUIDs, hash-ordered maps, absolute file paths, locale-sensitive number/date formatting, thread interleaving. You must inject a fixed clock/seed or scrub those fields before comparison, otherwise the suite becomes flaky and people start ignoring it. ## What to do when the behaviour looks wrong This is the most-tested judgement point in interviews. The rule: **the characterization test records reality, including bugs.** Reasons: - Callers may already depend on the quirk (Hyrum's Law: with enough users, every observable behaviour of your system will be depended on by somebody). - You cannot tell, from inside a refactoring, whether the quirk is load-bearing. - Mixing a behaviour change into a refactoring destroys the refactoring's safety property: when the suite goes red you no longer know whether you broke something or intentionally changed it. So: write the assertion capturing the odd value, add a comment or a distinctive test name marking it as suspected-incorrect, raise a ticket/ask the product owner, and if the fix is approved, do it as a **separate commit** where the changed characterization test is the visible diff. ## Trade-offs and limitations - **They are not documentation of intent.** A future reader sees `assertEquals(4210, result)` with no explanation of why 4210 is right. Over time, replace valuable ones with intention-revealing specification tests as you learn the domain. - **They can lock in the wrong thing.** If you characterize at too coarse a boundary (e.g. through the HTTP layer only), a refactoring can preserve the outer output while corrupting an internal invariant. - **They cost real time** on genuinely hostile code, but the cost is bounded to the change point rather than the whole system. - **Over-specification.** Asserting a full toString dump makes every cosmetic change fail. Match assertion granularity to what you intend to keep stable. ## Where they sit in the workflow In Feathers' Legacy Code Change Algorithm the order is: identify change points, find test points, break dependencies, **write characterization tests**, then make changes and refactor. They are the safety net that makes the last step non-scary.
- How is a characterization test different from a unit test written with TDD?Intent and failure semantics. A TDD test states required behaviour and is written red, before the code exists; it is a specification. A characterization test is written against existing code, is expected to pass immediately, and states only 'this is what happens today'. A TDD test failing means the code is wrong; a characterization test failing means the behaviour moved.
- You cannot even construct the class to characterize it — the constructor opens a database connection. What now?Break the dependency first with the smallest safe move: extract the connection creation into an overridable factory method and subclass it in the test, extract an interface and pass a fake in, or use a link/build-time seam to substitute the driver. Alternatively use Sprout Method so the new logic lives in a testable unit even if the host class stays untestable for now.
- How many characterization tests are enough?Enough to cover the behaviour your planned edit could disturb — typically driven by branch coverage of the change region plus its callers' observable effects. Full-system characterization is not the goal.
It is like taking a plaster cast of a machine part before you start filing it: the cast is not a blueprint of the ideal part, it is a record of the exact shape you had, so afterwards you can tell whether you removed anything you did not mean to.
saying these in an interview costs you the question
- Fixing an obviously-wrong value while writing the characterization test, in the same commit as a refactoring
- Calling characterization tests 'specifications' or treating them as requirement documentation
- Using random or arbitrary inputs instead of boundary- and branch-driven inputs
- Golden masters containing timestamps, UUIDs or unordered map output, then blaming flakiness on the tool
- Asserting only 'no exception thrown', which detects almost no behaviour change
- Deleting a characterization test because it 'asserts a bug'