skip to content

How do you check a generated test's expected values when the rule they encode is written down nowhere?

level: middleimportance: must knowfreq 55%

answer

  1. Two different checks hide in one question
  2. Some of it you can settle alone
  3. Concrete cases get answers, general ones do not
  4. The test becomes the written-down rule

basics

~20 s

Split it in two. Settle what you can alone - unit, scale, sign, which side of the boundary the value falls on - then take what is left to whoever owns the rule, as one concrete case with a yes-or-no answer.

solid answer

~50 s

Two different checks hide inside "is this value right". The first you can do alone: does the expectation have the right unit and scale, is its sign possible, does it fall on the side of the boundary the test's name implies, could the function even produce it? Those cost a minute and need nobody else's time. The second needs a person, because the rule lives in somebody's head and reading the code will not recover it - the code is one person's past reading of that rule, which is the thing in doubt. The move that works is to stop asking how the feature works and ask one concrete case: someone joins on the last day of the policy year, do they accrue that day? That gets answered. Then put the answer in the test's name, so the next reader inherits the rule and not just the number.

go deeper

for a junior

Know that you can check a lot on your own - unit, scale, sign, which side of the boundary - before you need anybody else, and that asking is the normal next step, not an admission of ignorance.

for a middle

Explain why reading the implementation cannot verify an expectation, and show how a concrete case turns an unanswerable question into one somebody can answer in a sentence.

for a senior

Demonstrate the whole loop including its failure: what you do when nobody can decide, how a pin gets labelled, and how the open question reaches the people who own the product.

for a principal

Own where rules get written down once they are settled, so the same question is not re-asked every time somebody generates tests against the same feature.

## Two questions wearing one coat "Is this expected value right?" is really two questions, and they have different answers, different costs and different owners. - **Is it a possible value?** Right unit, right scale, right sign, right side of the boundary the test claims to be about, reachable at all from the inputs the test supplies. This is arithmetic and you can do it alone, in about a minute per test. - **Is it the value the business wants?** This is a question about a rule, and no amount of reading the code answers it. The code is one person's past reading of that rule, on a day when they may have been wrong, and it is exactly the thing the test exists to check. Separating them is most of the skill. People skip straight to the second, find they cannot settle it, and then wave the whole test through. ## What you can settle on your own Do these in order, cheapest first, and stop as soon as one fails: 1. **Unit and scale.** Days where the rule says days, not hours; a whole number where the rule works in whole numbers. A value in the wrong unit is a defect in the test even if the code agrees with it. 2. **Sign and range.** Can this function produce a negative figure at all? Is the value larger than the maximum the rule allows anyone to hold? 3. **Boundary side.** A test named for service that has not completed a month should expect the answer for *not completed*. It is worth reading the name and the value together: an expectation can sit one step over the line its own name describes. 4. **Reachability.** Could these inputs produce that output under any reading of the rule? If not, the expectation came from a different case than the one the test sets up. 5. **Consistency across the file.** Two tests in the same batch that encode incompatible rules is a strong signal, and it costs nothing to notice. ## The question that actually gets answered For what is left, you need the person who owns the rule - a product owner, a specialist, the person who wrote the policy, occasionally a customer-facing colleague who has answered the question before. Ask them the wrong way and you get nothing. *"How does leave accrual work?"* invites a summary that omits the very edge you are asking about, or a promise to find out. Ask the concrete case and it is answerable in a sentence: *"Someone joins on the last day of the policy year and leaves on the first day of the next. Do they accrue anything for the day they joined?"* A yes or a no settles the expectation. Then the answer has to survive the conversation: - **Put the rule in the test's name**, not in a comment nobody reads: `a_joiner_on_the_last_day_of_the_policy_year_accrues_nothing`. - **Note who said so and when**, briefly, where your team keeps such things. The next person to see this test go red needs to know whether the rule is contested. - **Expect the conversation to find a defect.** If the owner's answer disagrees with what the code does, the generated test just did its job - decide the rule first, then decide whether the code or the test changes. ## When nobody can answer Sometimes the rule genuinely does not exist yet: the situation has never occurred, and nobody has decided. That is a finding about the product rather than about the test, and it should leave the building as one. Meanwhile the honest thing to do with the test is to keep it as a **pin on current behaviour** and to say so in its name, so that the next reader does not mistake a record of what the code does for a statement of what it should do. An unlabelled pin is worse than no test, because it will be cited as evidence. ## Why the model's own explanation is not the check It is tempting to ask the assistant where the number came from. The reply will be fluent and it will be generated from the same material that produced the number, so it cannot be independent evidence about the rule. Treat it as a hypothesis worth a minute - it often points at which line of the implementation the expectation was read off, which is useful - and never as the answer to the second question above. | | before the check | after the check | | --- | --- | --- | | the expectation | a number of unknown origin | a number derived from a stated rule, or a labelled pin | | the test's name | describes the inputs | states the rule | | what a failure means | something changed | the product's promise is broken, or the rule moved | | who can triage it | whoever wrote the prompt | anyone who can read the test |

  • The person who owns the rule disagrees with what the code does. What now?
    You have found a defect, which is the outcome you were hoping for. Decide the rule first and record it, then decide whether the code or the expectation changes - they are separate decisions and doing them in that order stops the code winning by default. Put the settled rule in the test's name so it does not have to be rediscovered.
  • How do you keep this from costing more than writing the tests by hand?
    Spend it where expectations are cheap to get wrong: the boundary cases, and any value you cannot derive in your head. The ordinary middle of the range rarely repays a conversation. One batched round of concrete questions usually covers a whole feature, and the answers stay useful long after that batch of tests.

saying these in an interview costs you the question

  • The code is the specification, so check the expected value against the code.
  • If the model explains where the number came from, that settles it.
  • If nobody wrote the rule down, there is nothing to verify against.
  • A plausible-looking number in a generated test does not need checking.
  • Rounding and units are details; the logic is what matters in review.