Why build policy fixtures from real merged changes instead of hand-writing the input?
answer
- the rule and the fixture share one author
- same wrong assumption on both sides
- shapes you would never type
- merged means the verdict is already decided
- captured corpus cannot contain the unusual
basics
~20 sA hand-written fixture encodes the author's mental picture of the input document — the same picture the rule encodes. If that picture is wrong, both are wrong together and the test still passes. Captured documents carry shapes nobody imagined.
solid answer
~50 sWhen I invent a fixture, I write down what I believe the input looks like. My rule reads the input the same way. A shared misunderstanding cancels out: the test is green, and the rule fails the first time it meets a real change. Capturing the plan documents from the last forty merged pull requests removes that circularity — they contain shapes I would never have typed, such as a resource created through a shared module, twenty resources in one change with one that matters, or an attribute set at a level I did not expect. Each captured case also arrives with its verdict already established: a human reviewed it and it merged, so the rule must allow it. That is ground truth rather than my opinion. Hand-written fixtures stay useful for the deny side and for branches no real change has produced yet.
go deeper
Know that fixtures are the recorded inputs a rule is tested against, and that ones taken from changes which already merged are more trustworthy than ones typed from memory.
Be ready to explain the circularity: the author's assumption about the document shape appears in both the rule and the invented fixture, so a wrong assumption cannot fail the test.
Demonstrate the operational side — trimming, freezing, redacting and labelling provenance — and be honest that a captured corpus is blind to shapes nobody has produced yet.
Own the standard: which rules must ship with evidence drawn from real merged work before they may be enforced, and who is accountable for keeping that evidence honest.
## The circularity problem Writing a rule and writing its fixture are the same act of imagination performed twice. Both start from the author's model of the document the engine will be handed. If that model is wrong — the attribute lives one level deeper than you thought, the field is a list where you assumed a string, the resource address is not the shape you expected — then the fixture is wrong in exactly the way that makes the rule look correct. The suite goes green and stays green until the rule meets production traffic, where it either misses everything or blocks everything. This is the specific reason "we have tests" is a weak claim for a policy rule. Tests over ordinary code compare behaviour against an expectation stated in a different language from the implementation. A fixture you typed yourself states the expectation in the same language, from the same head, at the same hour. ## What capture gives you For an infrastructure rule — say, only approved instance families in approved regions — the raw material is the plan document produced for changes that have already been merged. Take the last forty pull requests that touched instances and keep the plan JSON each one produced. Three things come with them for free. **Realistic shape.** Real changes are messier than invented ones. One resource is created inside a shared module and its attributes come from the module's variables, not from the caller. Another change creates nineteen unrelated resources plus one instance. A third uses a count or for_each so there are several addresses derived from one block. A fourth touches an instance without creating it. Each of those shapes is a way your rule can be wrong that you will not think of at a blank page. **A verdict you did not choose.** Every captured change carries an outcome that was decided by someone else: it was reviewed and it shipped. So the expected result is *allow*, and that expectation is a recorded fact rather than an assertion of your taste. If your candidate rule denies one of those forty, you have found either a real violation that slipped through review — worth knowing — or a defect in the rule. Either way, the fixture is doing work an invented one cannot. **A ready-made near-miss generator.** Once you hold a realistic allow fixture, the matching deny case is a one-field edit of it: change the region to an unapproved one and change nothing else. The pair now differs exactly at the boundary the rule draws, which is what makes the deny case meaningful rather than a strawman. ## How to keep the corpus usable - **Trim and freeze.** Store the smallest excerpt that still exercises the rule, and treat it as an immutable artifact checked in beside the rule. A fixture that is regenerated on every run is not a fixture. - **Label provenance.** Record, next to each case, whether it was captured or synthesised, and for captured ones which change it came from. A reviewer reading the table needs to know which rows are evidence and which are the author's construction. - **Scrub before you commit.** Real plans carry account identifiers, bucket names, internal hostnames and sometimes values that should not sit in a public repository. Capture is not an excuse to check in production detail wholesale. - **Keep them small in number.** Forty realistic cases beat four hundred; past a point you are re-testing the same shape and paying for it on every run. ## What capture does not give you A captured corpus is a sample of what teams have already done and had approved. It cannot contain a shape nobody has produced yet, and it cannot contain the changes that were rejected in review before they ever generated a merged plan. So it is systematically blind to two things: branches of the rule that only unusual inputs reach, and the deny side of the boundary. Both of those still need hand-written cases — the discipline is that you write them deliberately and mark them as constructed, rather than defaulting to invention for everything. ## Interview register The answer that lands names the circularity first, then the concrete shapes capture surfaces, then the ground-truth verdict, and finishes by conceding the blind spot instead of overselling the technique. An answer that says only "real data is better" has not shown why.
- Your rule denies one of the forty captured plans. What do you conclude?Not automatically that the rule is wrong. Read the change: either it really did violate the policy and review missed it, which is a finding worth reporting, or the rule is over-broad and this is a genuine near miss you now own as a permanent allow case. The value of the corpus is that it forces that conversation before the rule is enforced on anyone.
- How large should the captured corpus be?Large enough to cover the distinct shapes, not the distinct changes. Group the captured plans by structure — declared directly, produced by a module, many resources in one change, an update rather than a create — and keep a couple of each. Beyond that you pay run time to re-assert the same thing, and reviewers stop reading the table.
- What do you do about identifiers inside a captured plan?Redact before committing. Replace account numbers, hostnames and any credential-shaped value with stable placeholders, and keep whatever the rule actually reads intact. Redaction has to be consistent, because if a fixture's region or family field is scrubbed you have thrown away the very thing the case exists to exercise.
Marking your own exam with the answer key you wrote from memory. Everything agrees, and none of it is checked against the world.
saying these in an interview costs you the question
- Believes hand-written fixtures test the rule independently
- Captures the whole plan without trimming or redacting
- Regenerates fixtures on every run instead of freezing them
- Claims a captured corpus removes the need for synthetic cases
- Assumes a merged change is automatically compliant