In a generated change, why is a plausible call to a real API harder to catch than an invented one?
answer
- Two failures share one name
- A missing name announces itself
- Call and neighbours share one belief
- Check the contract at its source
- The change's own usage proves nothing
basics
~20 sA name that does not exist fails the moment anything resolves it. A real name used on a wrong assumption resolves and fits the code around it, since one run wrote both sides; the disagreement is with the real contract.
solid answer
~40 sTwo different things get called a hallucinated API. A name that does not exist has nothing to resolve to, so it fails as soon as anything tries — at build time where that is checked ahead of running, at first execution otherwise. A real name used on a wrong assumption about units, ordering, emptiness or mutation resolves and runs. It also reads as correct, because one run wrote the call and the code around it from the same belief, so the change agrees with itself; the disagreement is with the real contract, which lives outside the change. Catching it means reading the contract where it is defined instead of inferring it from how the change uses it, then naming the input that would make the assumption matter and checking whether anything exercises it.
code
pseudocode · 15 lines# file 3 of 9 - the helper the run wrote
function itemsInBinOrNone(bin):
return catalog.contentsOf(bin)
# file 7 of 9 - the caller, trusting that name two files away
function countSheetFor(bin):
items = itemsInBinOrNone(bin)
if items is empty:
return emptySheet(bin)
return sheetOf(items)
# the real contract: contentsOf reports an error for an unknown bin,
# and returns empty only for a bin that is known and empty. So the
# unknown-bin path above never runs, and the error leaves countSheetFor
# for a caller that has no handler for it.go deeper
Know that generated code can call something that does not exist, and that a call which builds cleanly may still be wrong about what the thing does. Check unfamiliar calls against their own documentation.
Explain why the plausible case is the dangerous one: the call and the code around it came from the same assumption, so the change agrees with itself and disagrees only with the real contract.
Show how you budget verification across a large change — which calls you check, why blast radius and silent failure outrank unfamiliarity, and what you say about the paths you did not verify.
Own the position that a review claiming to have verified every call in a large generated change has verified none. Make the coverage of a review explicit rather than implied.
## Two failures share one word "Hallucinated API" covers two failures that behave nothing alike in review. 1. **A name that does not exist.** The change calls something the codebase and its libraries simply do not have. 2. **A real name used on a wrong assumption.** Everything exists; the belief about what it does is wrong — the unit, the ordering, what the empty case returns, whether it copies or mutates, whether it reports failure by signalling or by returning a value, whether the result is ready or still pending, or whether the name belongs to a neighbouring type. The first is self-announcing. There is nothing for the name to resolve to, so it fails as soon as anything tries to resolve it: at build time where that is checked ahead of running, and at the first execution of the line where it is not. The failure usually names the symbol for you. **The second resolves, builds, runs, and reads beautifully.** ## Why the plausible one is coherent The mechanism is worth stating precisely, because it is the whole answer: **the call site and the code around it were written by the same run, from the same belief.** Suppose the run believed a bin lookup returns an empty result when the bin is unknown, where in fact it reports an error. Then the caller branches on empty, the helper above it is named as though empty means unknown, and the code below handles the empty case gracefully. Everything inside the change agrees with everything else inside the change. The disagreement is with the real contract, which lives outside the change — so reading the change alone, however slowly, will not surface it. That is also why "I would have noticed" is not a check. A reviewer cannot tell by reading which contracts the run could see and which it had to guess at. ## The assumptions that go wrong | the assumption | what a wrong one looks like at the call site | |---|---| | unit | a quantity treated as pallets where the source returns individual items | | absence | an empty result expected where the source reports an error, or the reverse | | ordering | results consumed as though sorted when the order is unspecified | | identity | a returned collection modified in place when it is a copy, or shared when it was assumed private | | timing | a result used immediately when it is not yet complete | | owner | a real name that belongs to a neighbouring type, not to this one | Each of these produces code that builds. Each fails only on the input that makes the assumption matter. ## Why a multi-file change hides it better In a change of one file you hold the assumption and its use in your head at once. Across nine, the belief is often **stated in one file and relied on in another** — a helper wraps the call and its name encodes the belief; a caller two files away trusts the name. Read either file alone and it reads correctly. Nothing is wrong in either place; the wrongness is in the relationship, and the relationship is what a file-by-file read never looks at. Attention makes it worse: it goes to the files that carry the feature, while the quiet wrapper looks like plumbing. ## How to actually check 1. **Choose the calls worth checking.** Unfamiliar names, anything crossing a boundary you do not own, anything whose result is used without a branch, anything that writes data. 2. **Read the contract where it is defined or documented** — not the change's own usage of it. **The change's usage is the thing under test; it cannot also be the evidence that the usage is right.** 3. **Name the input that would make the assumption matter**: the unknown bin, the empty count, the second call in a row. 4. **Find out whether anything exercises that input today.** Where nothing does, either exercise it or say plainly in the review that this path is unverified. ## Where to spend the attention You cannot verify every call in a nine-file change, and a review that claims to have done so has verified nothing. Rank by blast radius rather than by suspicion: calls that write or delete data, calls across a boundary you do not own, and calls whose wrongness would be **silent** rather than loud. A wrong assumption that crashes is a bad afternoon; a wrong assumption that quietly writes the wrong stock level is a bad quarter. ## What this does not mean - **Not every unfamiliar call is a hallucination.** Most are simply unfamiliar to you, which is a fact about the reader rather than about the code. - **The failure is not unique to generated code.** A person writing against an unfamiliar library makes exactly the same mistake. - **What differs is volume and presentation.** A run produces call sites faster than anyone reads them, and the code around a guessed contract rarely carries the hedging — a tentative comment, a question raised before merging — that a person unsure of a contract tends to leave behind.
- Is a hallucinated API a different problem from a wrong assumption about a real one?In review, yes. A name that does not exist fails the first time anything resolves it, so it costs you a build rather than an incident. A real name used wrongly resolves, runs, and fails only on the input that makes the assumption matter — which may be rare and may be expensive.
- The change wraps the suspicious call in a helper. Does that help or hurt?Both. A wrapper concentrates the assumption in one place, which is where you want it. It also states the assumption in its own name, so reading the callers raises no question at all. Check the wrapper against the real contract rather than against the way its callers use it.
- Which calls do you skip when a change is too large to verify exhaustively?Skip the familiar ones you use daily and the ones whose wrongness would be loud and immediate. Spend the attention where a wrong assumption would be silent: calls that write data, and calls across a boundary you do not own. Then say in the review which paths you did not verify.
A signpost to a town that does not exist is caught the first time anyone looks at a map. A real town name on a sign pointing the wrong way survives every check that only verifies the name.
saying these in an interview costs you the question
- If it builds, the call must be using the function correctly
- Hallucinations are always invented names that fail to resolve
- The surrounding code confirms what the function does
- Ask the agent whether the call is right and accept the answer
- Generated calls are fine as long as the names are spelled correctly