What is a test oracle, and what makes one fallible?
answer
- Something has to say it is wrong
- An observation is not yet a verdict
- Kinds differ in strength and cost
- Each one is a model, not truth
- Shared origin means shared blind spot
basics
~20 sA test oracle is whatever you use to decide that an observed behaviour is wrong: a written specification, a standard, a comparable product, an earlier version, or your own expectation. Each is an imperfect model, so it can be wrong too.
solid answer
~50 sA test has two halves: the procedure, which produces an observation, and the oracle, which turns that observation into a verdict. Oracles come in very different strengths - an independently computed exact value, a written requirement, a published standard, the product's own documented claims, a comparable product, the previous release, the product's consistency with itself, a tester's expectation of what a reasonable user would want, and at the weak end the implicit oracle that it did not crash, hang or corrupt data. All of them are heuristic: the requirement can be stale, the previous release can carry the same defect, expectation can be uninformed. So an oracle disagreement means *there is probably a problem here worth reporting*, not *the code is definitely broken* - and a bug report is stronger when it names which oracle was contradicted.
code
pseudocode · 7 linesobservation = run_case(input)
# the oracle is the judging half, and it has a source
if observation != statutory_rounding_rule.expected_for(input):
report(problem, oracle = "published rounding rule", confidence = "high")
else if observation != previous_release.result_for(input):
report(question, oracle = "previous release", confidence = "low, may enshrine an old defect")go deeper
Be ready to define an oracle in one sentence and list four or five kinds. The point interviewers listen for is that the oracle is the deciding half of a test, separate from the steps you performed.
Explain how the kinds differ in strength and cost, and give a concrete failure mode for two of them - a stale requirement, a previous release that carries the defect you are hunting.
Show that you notice which oracle you are leaning on during real investigation, and that you name it in findings. Demonstrate the two error directions, false alarm and miss, and why misses cluster where oracle and implementation share an origin.
Own the argument that oracle independence, not test count, is what determines whether a suite can catch a misunderstanding, and be able to say where a team should spend to obtain a stronger, independent one.
## The two halves of a test Every test does two separable things. It **exercises** the product - drives inputs, walks a workflow, feeds a data set - and it **judges** what came back. The judging half is the oracle. In the testing literature an oracle is defined as a mechanism for *recognising a problem*, and that phrasing is deliberate: oracles are much better at saying 'this is wrong' than at saying 'this is right'. You can exercise a system beautifully and learn nothing if you have no basis for deciding that what you saw was a failure. This asymmetry is why the oracle, not the input, is usually the hard part. Choosing inputs is a design activity with well-known techniques. Deciding that a number on a screen is the *wrong* number, when nothing written down says what the right one is, is the oracle problem, and it is where most exploratory testing effort actually goes. ## Kinds of oracle, roughly by strength - **An independently derived exact value.** Someone computes the answer by another route and compares. Strongest and usually most expensive. - **A specification or requirement.** Strong when current and precise; silent on everything it does not mention. - **A standard, statute or regulation.** Strong and externally defensible: rounding rules, date formats, accessibility rules, tax rules. - **The product's own claims.** Help text, release notes, on-screen labels, the sales page. Cheap, and a contradiction is hard to argue with. - **A comparable product.** Something else that solves the same problem. Suggestive, never binding - the other product is allowed to be different, and allowed to be wrong. - **The previous version.** The regression oracle. Cheap and high-volume, but it enshrines whatever the previous version already got wrong. - **Internal consistency.** The same fact shown in two places must agree; a feature must behave like its siblings. - **User expectation and purpose.** What a reasonable user of this thing would want, and whether the product serves the job it exists to do. - **Implicit oracles.** No crash, no hang, no unhandled error, no corrupted data, no runaway resource use. They apply to every test for free and catch only catastrophes. ## Why every oracle is heuristic An oracle is a *model* of correct behaviour, and a model is a simplification that is right most of the time and wrong sometimes. Each kind fails in its own way. A requirement document drifts out of date faster than the code. A comparable product implements a different, equally defensible policy. The previous release is the ancestor of the defect you are hunting. Expectation reflects a tester who is not the user. Oracle error runs in two directions, and they cost very differently: - **False alarm** - the oracle says wrong, the product is right. Costs investigation time and, if it repeats, credibility. - **Miss** - the oracle agrees with a wrong product. Costs a shipped defect. Misses cluster wherever the oracle and the product share an origin: a requirement and an implementation written from the same misunderstanding will agree perfectly with each other and both be wrong. **Independence is what gives an oracle its power.** ## What that means in practice A worked example. A payroll engine is rewritten, and in month-end run 47 the net-pay totals differ from the ledger export by fractions of a currency unit - a drift of about 0.02 per payslip across 3,412 payslips. Which oracles are available? The requirement document does not state a rounding rule. The previous release rounds differently, but it is the thing being replaced, so it cannot arbitrate. The statutory guidance for the jurisdiction does state a rule - that is the strong, independent oracle, and it settles the case. Internal consistency already proved *something* is wrong, because the payslip total and the ledger export are two views of one number and they disagree; it simply could not say which view was at fault. That is the everyday shape of the skill: notice that you are relying on an oracle, name which one, know how it can fail, and reach for an independent one when the stakes justify the cost. Say it in the report, too - 'this contradicts the published rounding rule in the statutory guidance' is a finding a developer acts on; 'this looks wrong to me' is an argument.
- If a test makes no explicit check at all, does it have an oracle?Yes - a very weak one. Simply running to completion exercises the implicit oracles: it did not crash, hang, throw an unhandled error or corrupt its data. Those catch catastrophes and nothing else, so such a test can pass while the output is entirely wrong. Say that plainly rather than treating the clean run as evidence of correctness.
- What is the difference between an oracle and an expected result?An expected result is one concrete value for one case. An oracle is the source or rule that produced it, and it can judge cases nobody enumerated - a rule such as 'net pay never exceeds gross pay' renders a verdict on every run there will ever be. Expected results are instances; the oracle is the generator and the justification behind them.
- Why does it matter whether an oracle is independent of the implementation?Because an oracle that shares an origin with the code shares its mistakes. If the same person derived both from the same misread rule, they will agree with each other and both be wrong, and the test will pass forever. Independence - a different derivation, a different author, an external standard - is what lets an oracle catch a misunderstanding rather than only a typo.
An oracle is a referee, not a physics engine. It makes the best call available from the evidence it has, and a call being occasionally wrong does not make the match unrefereeable.
saying these in an interview costs you the question
- Claims only a written specification can serve as an oracle
- Treats a passing check as proof the behaviour is correct
- Assumes the previous release is by definition right
- Reports something as wrong without naming what it contradicts
- Confuses the oracle with the inputs chosen for the test
- Believes a strong oracle removes the need for judgement