How do you decide whether a stronger, costlier oracle is worth building?
answer
- It is an investment, not a checklist
- What would a wrong answer cost?
- Silent failures justify expensive judgement
- Same source means shared blind spot
- Noisy oracles decay to zero strength
basics
~20 sWeigh what a wrong answer costs and how silently it fails against the oracle's build, run and maintenance cost, its false-alarm rate, and above all its independence from the implementation. Buy strength where failures are expensive and quiet.
solid answer
~50 sTreat oracles as a portfolio rather than a checklist. Four factors decide funding. **Strength** - how much wrongness this oracle could detect, not how many cases it covers. **Independence** - an oracle derived from the same source as the code shares its misunderstandings and is worth far less than its cost suggests. **Total cost** - build, run, and the maintenance that intentional behaviour change will impose on it forever. **False-alarm cost** - a noisy oracle is eventually ignored, which makes its effective strength zero. Then weight by consequence: buy a strong independent oracle where a wrong answer is expensive and **silent**, such as money, safety or a regulated calculation; accept weak implicit oracles where failures are loud, cheap and quickly noticed. The two anti-patterns to name are re-implementing the system as its own oracle, and freezing today's behaviour as the standard.
go deeper
The takeaway is that stronger ways of judging results cost more, so teams do not use the strongest one everywhere. Notice which one your team relies on in the area you work in.
Be able to compare two candidate bases for judging the same behaviour on strength and cost, and explain why one that never disagrees with the code may be catching nothing at all.
Argue a specific case with numbers: what the wrong answer costs, whether it fails loudly or silently, what the oracle costs to maintain through intentional change, and what its false-alarm rate does to its real value.
Own the portfolio and the accepted gaps. State the funding rule in advance, assign a maintenance owner to every expensive oracle, and be willing to say plainly which risks the organisation runs with only weak judgement in place.
## The decision, stated properly A lead is rarely asked 'is this oracle good?' The real question is where to spend a fixed budget of engineering time and expert attention so that the wrongness that would hurt most is the wrongness most likely to be caught. That reframes oracle choice as an investment decision with four inputs. **Strength.** How much wrongness could this oracle detect if it were present? An implicit oracle catches catastrophes only. A previous-release comparison catches unintended change. An independently derived value catches an incorrect answer that has always been incorrect. Strength is about the *class* of error detectable, not the number of cases executed - a thousand cases behind a weak oracle detect one class of problem a thousand times. **Independence.** This is the factor most often left out of the analysis and the one that most often decides the outcome. An oracle derived from the same requirement, by the same person, at the same time as the implementation shares its blind spots exactly. It will agree with the code forever and catch nothing but transcription slips. Where the risk is a misunderstanding rather than a mistake, only an independent oracle - an external standard, a separately derived calculation, a domain expert who was not in the design conversation - has any power at all. **Total cost of ownership.** Build cost is the visible part and usually the smaller one. The running cost matters if the oracle needs a person. The maintenance cost is the killer: every intentional behaviour change obliges someone to update the oracle, and a team that resents that obligation will weaken the oracle rather than maintain it. Ask how often the behaviour under judgement is expected to change on purpose. High-churn areas punish expensive oracles. **False-alarm cost.** An oracle that cries wolf gets ignored, and an ignored oracle has an effective strength of zero however strong it is on paper. Budget the false-alarm rate explicitly, and treat a rise in it as a defect in the oracle. ## Weighting by consequence Spend where a wrong answer is **expensive and silent**. Silence is the crucial half: a crash announces itself and a free implicit oracle catches it, so paying for strength there buys little. A slightly wrong number that nobody notices for a year is the case that justifies an expensive, independent oracle - regulated calculations, money movement, safety, anything that produces an output nobody can eyeball for plausibility. A worked decision. A payroll engine is being replaced by an 11-person team. Proposal: fund an independently derived recalculation of net pay, written from the published statutory rules by someone who did not build the engine, at an estimated 14 engineering days plus about 2 days of the domain expert's time. Against: the engine is money-moving, the failure mode observed in trials is a currency-rounding drift of roughly 0.02 per payslip that no reviewer would ever spot by eye, and the affected population is 3,412 employees per period. The alternatives - comparing against the outgoing engine, or cross-checking the ledger export - are cheap but dependent or partial: the outgoing engine may embody the same misreading of the rounding rule, and the cross-view check proves only that two figures disagree. On those facts, strength and independence justify the 14 days. Change one input - a low-value internal report with a human reader who would spot nonsense instantly - and the same 14 days are indefensible. ## Two anti-patterns worth naming **Re-implementing the system as its own oracle.** Writing a second implementation of the same logic from the same understanding gives you two things that agree and one budget spent twice. It has value only when the second derivation is genuinely independent - different source rules, different author, different route to the answer. **Freezing current behaviour as the standard.** Adopting the existing release as the authority is cheap, scales to enormous case counts, and quietly declares that everything the product does today is correct. It is a fine change-detector and a bad correctness oracle, and the failure is invisible: the suite is green precisely because it is measuring agreement with the thing you are trying to check. ## What a principal answer includes A decision rule stated in advance rather than argued case by case; an explicit statement of which risks the organisation accepts having no strong oracle for; a maintenance owner for every expensive oracle funded; and a review trigger - if the false-alarm rate climbs, or the behaviour enters a period of intentional churn, the investment is revisited. Interviewers are listening for someone who can say 'we deliberately do not have a strong oracle here, and this is the exposure that buys us' without embarrassment. Naming an accepted gap is a stronger answer than claiming full coverage.
- Why can a second implementation of the same logic be a worthless oracle?Because if it is derived from the same rules by the same understanding, it reproduces the same misinterpretation and agrees with the original forever. It then catches only transcription slips while costing a full second build. It becomes worthwhile only when the derivation is genuinely independent - different source material, a different author, a different route to the answer.
- How does an oracle's maintenance cost change your decision in a high-churn area?It usually decides it. Every intentional behaviour change forces an update, so in a fast-changing area an expensive oracle either consumes continuous effort or gets quietly weakened until it stops detecting anything. Prefer cheap, coarse oracles where change is frequent, and reserve strong expensive ones for stable, high-consequence logic that rarely changes on purpose.
- What do you do about an area you have consciously decided not to fund a strong oracle for?Record it as an accepted gap with the exposure named, not as covered. State what class of wrongness would go undetected, who accepted that, and what would trigger revisiting it - a regulatory change, an incident, or the area becoming money-moving. An explicit accepted gap is defensible; an unexamined one becomes an incident postmortem finding.
saying these in an interview costs you the question
- Judges oracle value by number of cases covered
- Ignores independence when comparing candidate oracles
- Counts build cost but never maintenance cost
- Treats a noisy oracle as still fully effective
- Adopts current behaviour as the definition of correct
- Claims full coverage rather than naming accepted gaps