Some teams record real API responses into fixture files and replay them in frontend tests; others hand-write the smallest stub the component needs. What does each approach buy and cost, and how would you choose?
answer
- fidelity versus legibility
- nobody reviews a huge JSON diff
- captured traffic carries real user data
- minimal stubs encode the happy shape
- record once, then reduce to a factory
basics
~20 sRecorded responses give real fidelity but are large, unreviewable, may carry secrets, and rot silently. Minimal hand-written stubs are readable and intent-revealing but encode only what the author believed. Most teams record once to learn the shape, then reduce it to a typed factory.
solid answer
~50 sRecording buys fidelity you cannot invent: the fields you did not know existed, the real nullability, the actual sizes and encodings. It costs review — nobody reads a 4,000-line JSON file — and it drags in timestamps, ids and sometimes real user data or tokens, plus it goes stale with no signal. A minimal hand-written stub is the opposite: readable, obviously tied to what the test asserts, but it is a transcription of one developer's belief and it makes tests pass on data the API never sends. I usually record once to discover the true shape, then reduce it to a typed factory whose defaults came from that recording, keeping one full recorded sample per endpoint as a schema-validated reference. I choose recording when the payload is large or poorly documented, and hand-written stubs when the component consumes a handful of fields.
go deeper
Be able to state the basic tradeoff: recorded payloads are realistic but huge, hand-written stubs are readable but only contain what the author thought of. Know that recorded traffic may contain real user data.
Explain the concrete failure modes of each — unreviewable diffs and embedded timestamps on one side, happy-shape-only data and invented fields on the other — and describe the record-then-reduce hybrid.
Show judgment about which realism actually buys detection: nullability, collection size, encodings and id formats, versus fidelity in fields nothing reads. Describe how you keep a recorded sample fresh and scrubbed rather than trusting it forever.
Own the policy: whether recording from production is permitted at all, what the scrubbing and retention rules are, and how fixture refresh is funded so realism does not decay into a false sense of coverage.
## Two ways to obtain a payload Every stub body comes from one of two places: a real response captured from a running service, or a developer typing what they think the response looks like. The tradeoff between them is fidelity versus legibility, and both extremes fail in their own way. ## What recording buys - **Fields you would never invent.** Real payloads carry metadata, links, audit fields and deeply optional branches that no one transcribes from memory. - **True nullability and formats.** The recording shows that `middleName` is genuinely `null` for most users and that `createdAt` really is an offset-bearing timestamp — details specs are frequently wrong about. - **Real scale.** A list endpoint that returns 200 items with 40 fields each exercises rendering and mapping code that a three-item stub never touches. - **Speed of authoring.** Capturing beats transcribing, especially for a poorly documented endpoint. ## What recording costs - **Unreviewable diffs.** A pull request that changes a 4,000-line fixture gets rubber-stamped. Nobody can tell an intentional change from an accidental one. - **Data hygiene.** Captured traffic can contain real names, emails, tokens and session identifiers. Recording pipelines need scrubbing, and a scrubber is code that can be wrong in a way that ends up in a public repository. - **Nondeterminism baked in.** Absolute dates, relative age calculations and ids captured from one environment make tests that pass today and fail next quarter. - **Silent staleness.** A recording made in March is indistinguishable from one made yesterday. Without a re-record job or schema validation, it rots exactly like a hand-written stub, only with more confidence attached. - **Test intent disappears.** When the fixture is enormous, the reader cannot tell which of its 200 fields the assertion depends on. ## What minimal hand-written stubs buy and cost The minimal stub is a communication device: the fields present are, by construction, the fields the test is about. It reviews well, it composes with factories, and it makes a test's dependencies explicit. Its weakness is that it is a belief. It typically contains the happy shape only — every optional field populated, every nullable non-null, every collection small. Code that would break on a real payload passes comfortably. It also tends to be *too permissive in the other direction*: developers stub fields the API does not actually send, and the component starts depending on data that will never arrive. ## The hybrid most mature teams land on 1. **Record once** against a real environment to discover the honest shape. 2. **Derive a typed factory** whose defaults are the realistic values from that recording, so tests read as `makeUser({ isAdmin: true })` rather than as walls of JSON. 3. **Keep one full recorded sample per endpoint**, scrubbed, checked into the repo, and validated against the schema in a conformance test. It documents reality and gives you something to re-derive from. 4. **Refresh on a schedule**, not on memory — a periodic job that re-records and fails when the new capture differs structurally from the stored one. That structure gets the fidelity of recording with the legibility of hand-writing, and it gives drift somewhere to show up. ## Realism that actually matters When deciding how faithful a fixture must be, the properties that repay effort are the ones the app's code branches on: nullable fields actually being null sometimes, collections large enough to hit pagination and virtualisation paths, strings with non-ASCII characters and lengths that break layout, and identifiers in their real format. Realism in fields nobody reads is pure maintenance cost. ## How to choose Record when the payload is large, poorly documented, or drives derived logic you cannot fully predict. Hand-write when the component consumes a handful of fields and the value of a readable test outweighs the value of unknown extra fields. In both cases the fixture must be typed from the schema and validated against it — that requirement is what makes either choice safe, and it is more important than which one you pick.
- What makes a recorded fixture nondeterministic, and how do you handle it?Absolute timestamps, ids and anything the app renders relative to now — a recorded `createdAt` becomes "3 months ago" and then "2 years ago". Either normalise those fields during capture to fixed values, or freeze the clock in the test so relative rendering is stable regardless of when the suite runs.
- Which properties of a payload are worth making realistic, and which are not?The ones the code branches on: real nullability, collection sizes that reach pagination or virtualisation paths, non-ASCII strings, and identifiers in their true format. Fidelity in fields no component reads is maintenance cost with no detection value.
- Why is a stub that includes fields the API never sends a real hazard, not just clutter?Because the component can start reading them. The test proves a feature works, production never supplies the field, and the failure appears only for real users. Typing the fixture against the generated schema type catches exactly this, since the extra property fails to compile.
saying these in an interview costs you the question
- Committing raw recorded responses including tokens or personal data
- Assuming a recorded fixture stays accurate indefinitely
- Calling a three-field happy-path stub realistic
- Keeping thousand-line fixtures nobody can review
- Recording payloads with live timestamps and current dates