How do you proceed when a result looks wrong but no oracle can settle it?
answer
- Re-running it will not add certainty
- List what could arbitrate, then rank
- Independence matters more than convenience
- Contradiction is reportable without a value
- The missing basis is itself a finding
basics
~20 sName it as an oracle problem: rank candidate oracles by strength and independence, prove a contradiction with a partial check such as cross-view consistency, spend expert time on one narrow question, and report the ambiguity if nothing resolves it.
solid answer
~50 sWork the oracle problem explicitly rather than staring at the output. First list every candidate: an external standard, the product's claims, cross-feature consistency, the previous release, a comparable product, purpose. Rank them by strength and, more importantly, by **independence** from whoever built the thing. Second, look for a *partial* oracle - bounds, totals and cross-view agreements that must hold - because a proven contradiction is a report even when nobody can name the correct value. Third, treat the human expert as a scarce oracle: arrive with one crisp question and the evidence, not with the whole investigation. Fourth, if nothing resolves it, the missing oracle *is* the finding: raise the ambiguity, state what would settle it, and note the exposure if the current behaviour is wrong. Never silently downgrade an unresolved observation to 'works as designed'.
code
pseudocode · 9 lines# a partial oracle: proves a contradiction without knowing the right value
payslip_total = sum(payslip.net for payslip in period_47)
ledger_total = ledger_export.total_for(period_47)
ytd_total = year_to_date.delta_for(period_47)
if not (payslip_total == ledger_total == ytd_total):
report("three views of one figure disagree",
oracle = "internal consistency",
note = "proves a defect; does not say which view is wrong")go deeper
Take away one rule: not knowing the correct value does not mean you must stay silent. Record what you saw and what you compared it against, and hand it to someone who can weigh it.
Be able to list candidate sources of judgement and pick between them, and to show that a cross-view contradiction is reportable even when nobody can name the right answer.
Demonstrate the whole sequence under time pressure - rank by independence, prove what a partial check can prove, spend expert time on one crisp question, and write up an unresolved case with its exposure stated.
Own the stopping rule and the escalation path: how much unresolved ambiguity the organisation tolerates on a money-moving path, and who is accountable for obtaining the missing authority.
## Why this is the hard case Most testing advice assumes you know what the right answer is. The genuinely difficult situations are the ones where you can see something suspicious and nothing available can arbitrate: an undocumented calculation, a legacy behaviour whose original author left, a numeric result that is plausible but not obviously right. Handling this well is a senior skill because it is not a technique, it is a sequence of judgements about evidence and cost. ## A worked incident A payroll engine is being replaced. In month-end run 47 the new engine produces net-pay totals that drift from the outgoing engine by fractions of a currency unit - about 0.02 per payslip, across 3,412 payslips, roughly 68 currency units for the period. Which is right? The requirement document is silent on rounding. The outgoing engine is the thing being retired, so its behaviour is a candidate, not an authority. The team is 11 people and exactly one of them, the payroll domain expert, is likely to know the actual rule - and she is the constraint on the whole migration. ## The sequence **1. Say out loud that this is an oracle problem.** The failure mode here is spending hours re-running the case, hoping certainty will arrive from repetition. It will not; nothing about the observation is in doubt. What is missing is a basis for judging it. **2. Enumerate candidate oracles and rank them.** Write the list: an external rounding standard for the jurisdiction; the product's own claims in the payslip footer and help text; internal consistency between payslip, ledger export and the year-to-date summary; the outgoing engine; a comparable product; the stated purpose of the calculation. Rank by two axes - how much wrongness the oracle could detect, and how independent it is of whoever built the new engine. An external published rule wins on both; the outgoing engine is cheap but dependent on the same lineage of assumptions. **3. Reach for a partial oracle first.** A partial oracle cannot name the correct value but can prove a contradiction. Bounds: net pay can never exceed gross. Totals: the sum of the payslip lines must equal the ledger export for the same period. Cross-view: three presentations of one number must agree. If any of those breaks, you have a defect without ever resolving the rounding question - and that report can be written today. Building a *machine-checkable* oracle is a distinct discipline with its own techniques; the point here is what you do while you do not have one. **4. Buy the human oracle efficiently.** On an 11-person team the domain expert is the strongest available oracle and the scarcest. Do not hand her the investigation. Arrive with one question that only she can answer - 'is the rounding applied per line or once on the period total?' - plus the two candidate outputs and the size of the difference. Ten minutes of her time is worth more than a day of yours, and framing it that tightly is what gets the ten minutes. **5. If nothing settles it, the ambiguity is the finding.** Raise it as an open question rather than silently letting it pass or inflating it into a defect you cannot support. A useful write-up states: what was observed, which oracles were consulted and why each was inconclusive, what evidence would settle it, and the exposure if the current behaviour turns out to be wrong. That last part is what makes it actionable for someone deciding whether to ship. **6. Set a stopping rule before you start.** Oracle hunting is unbounded. Decide up front how much of the timebox this one observation gets, on the basis of the risk if it is real. A fractional drift on a money-moving path earns more; a cosmetic ordering difference earns much less. ## What weak answers sound like - 'No spec, so it is not a bug.' This converts an unanswered question into a false negative, and it is the single most damaging habit in this area. - 'I asked the developer.' The developer is an oracle, but the least independent one available - if the behaviour comes from a misunderstanding, they hold that misunderstanding. Use them for mechanism, not for verdicts. - 'I filed it as high severity to force an answer.' Inflating impact to buy attention burns the credibility you need next time. - 'I compared with a competitor and filed the difference.' A comparable product raises a question; it does not establish a requirement. The underlying stance is simple and it is what interviewers are listening for: an unresolved observation is information, and information is reported. You are not required to know the right answer to be entitled to say that something looks wrong - you are required to say what you compared it against, and how confident that comparison makes you.
- Why is the developer a poor oracle for a suspected behaviour, and when are they the right person to ask?They are the least independent source available: if the behaviour comes from a misunderstanding, they share it, and asking them frequently converts a real defect into 'works as designed'. Ask them about mechanism - what the code does and why it does it - and take the verdict from something independent, such as a published rule or a domain expert who was not involved in building it.
- How do you write up an observation you could not resolve, without inflating or burying it?State the observation precisely, list the oracles you consulted and why each was inconclusive, name the evidence that would settle it, and quantify the exposure if the behaviour is wrong. Log it as an open question rather than a confirmed defect. That gives a decision-maker exactly what they need and preserves your credibility for the findings you can prove.
- How much time should one unresolved observation get?Decide before you start, in proportion to the risk if it is real. A fractional discrepancy on a money-moving path justifies chasing an external rule and interrupting a domain expert; a cosmetic difference gets a note and a move on. Without a stopping rule, oracle hunting expands to fill the whole session and the rest of the charter goes unexamined.
saying these in an interview costs you the question
- Says nothing can be reported without a specification
- Re-runs the same case hoping certainty will appear
- Takes the implementer's word as the verdict
- Inflates severity to force somebody to answer
- Hands the domain expert the whole investigation
- Lets an unresolved observation quietly disappear