Automated scanners reliably report injection findings and outdated-dependency findings, but they mostly miss the case where an authenticated request returns a record belonging to a different customer — which is exactly the shape the OWASP Top 10 ranks first by incidence. Explain the testing-theory property that separates those categories, and what it implies about how the Top 10's own numbers should be read.
answer
- oracle = decide wrongness from the observation alone
- injection is self-evidencing; a 200 is not
- authorised and unauthorised responses are byte-identical
- missing authorization: absence has no location to review
- incidence = detected prevalence, a floor; survey slots patch the blindness
basics
~20 sTesting needs an oracle: a way to judge wrongness from the observation alone. Injection and vulnerable components have one. An authenticated 200 carrying someone else's record does not — ownership is a domain fact absent from the artefact under test.
solid answer
~50 sAn automated finding needs an **oracle** — a rule that decides, from the observation alone, that the behaviour is wrong. Injection has a strong one: input crossing from data into an interpreter's grammar is wrong in every application in every domain. A dependency at a version with a published advisory is decidable from an inventory. A missing transport or cookie directive is decidable from a response header. An access-control failure produces a well-formed, authenticated, correctly-routed 200. The bytes of an authorised response and an unauthorised one are identical; only a **domain fact** — who owns this record, which tenant, which workflow state — separates them, and that fact appears nowhere in the traffic or the code. No oracle, so no tool can decide it; the oracle has to be supplied from outside, as an explicit policy model. So the Top 10's incidence percentages measure *detected* prevalence, biased toward decidable shapes. That bias is why the list reserves slots filled by practitioner survey rather than by contributed data.
code
text · 8 linesREQUEST GET /orders/1042 Cookie: session=<valid>
RESPONSE 200 OK {"id":1042,"total":89.90,...}
case A: order 1042 belongs to the session principal -> feature
case B: order 1042 belongs to a different customer -> breach
case C: principal is a support agent with a read grant -> feature
observable difference between A, B and C: nonego deeper
Be able to say what an oracle is in plain terms — a way to tell that an output is wrong — and give one category that has one (injection, outdated dependency) and one that does not (a normal-looking response carrying someone else's data).
Explain why the two responses are indistinguishable on the wire, and add the absence point: a missing check has no line number, so reviewers and analysers that examine what exists cannot see it.
Draw the conclusion for assurance planning: tooling coverage is a function of oracle strength, not of risk, so a clean scan says nothing about this category, and its reported incidence is a floor.
Frame the Top 10 as an instrument with known bias and reason about what it therefore cannot tell you about your portfolio — including why the list reserves survey slots, and why programme-level assurance investment should not be allocated in proportion to what the scanners report.
## What an oracle is In testing theory, an **oracle** is whatever lets you decide that an observed behaviour is wrong. Running the system produces an observation; the oracle turns that observation into a verdict. Without one you can execute a program forever and learn nothing, because you have no basis for calling any output a defect. Security categories differ enormously in how strong their oracle is, and that single property — not how serious the bug is, not how common it is — determines whether tooling can find it. ## Categories with an intrinsic oracle Some defect classes carry their verdict with them. - **Injection.** The invariant is universal: attacker-supplied data must never be lexed as part of an interpreter's grammar. If a probe changes the *structure* of a query, a command line or a document rather than only its values, that is wrong for every application in every domain. The observation is self-evidencing — no knowledge of the business is needed. - **Vulnerable and outdated components.** Decidable from an inventory plus a public advisory database. A version string either falls in an affected range or it does not. - **Security misconfiguration** in its mechanical forms — a missing cookie attribute, a permissive cross-origin policy, a debug endpoint, a default credential. The response or the config file is the evidence, and the rule is application-independent. These are the categories where a scanner is genuinely useful, and unsurprisingly they are the categories whose findings dominate automated reports. ## Why broken access control has no intrinsic oracle Now take the request that returns another customer's order. It was authenticated. It was routed to the correct handler. It executed a syntactically correct query. It returned a schema-valid body with status 200. Every layer behaved exactly as written. The only thing wrong is that *this* principal should not have received *that* resource — and "should not" is a statement about ownership, tenancy, role, delegation, purchase or workflow state. None of that appears in the HTTP exchange, and none of it is derivable from the code, because the code is precisely what omitted it. Two byte-identical exchanges can be a legitimate feature in one application (a support agent reading any order) and a breach in another. The verdict lives in the domain, not the artefact. A useful contrast: a spellchecker flags *teh* without knowing what you meant, but cannot flag *form* written where you meant *from*. Both are valid words; only intent decides. Access control is the second kind of error. ## A defect of absence There is a second, compounding property. Injection is something you **wrote** — a concatenation exists somewhere, at a line number. Broken access control is usually something you **did not write**: the check that was never added to a new endpoint, the ownership predicate missing from one of four read paths, the bulk export that bypasses the guarded handler. The CWE naming captures it directly — *missing authorization* is a distinct weakness from *incorrect authorization*. Missing code has no location. Reviewers and static analysers both work by examining artefacts that exist; an omission presents no artefact to examine. Combine that with ordinary software evolution — an endpoint duplicated for a mobile client, a report bolted onto a legacy service — and omission becomes the steady-state failure mode rather than an unlucky one. ## What the incidence number actually measures Here is the consequence that matters for reading the taxonomy. Not all access-control failures are undecidable: a completely unauthenticated endpoint returning data, a directory listing, a token accepted with its signature unverified — these have oracles, because *no* domain rule permits them. Tools do find those. So the measured incidence for the category is drawn largely from its decidable subset plus whatever human-assisted testing contributed. Which means the number is a floor, not an estimate. The category tops the incidence tables *despite* the measurement being partly blind to its dominant shape, which makes the true rate higher than reported, not lower. Anyone who reads Top 10 percentages as prevalence has inverted the argument. ## The consequence for the list itself The Top 10 is compiled from contributed testing data across a large application population, and incidence is defined as the share of tested applications with at least one instance of a mapped weakness. That method inherits the bias described above: it can only count what current testing detects. The list's authors handle this explicitly by filling a minority of the slots from a practitioner survey instead of from the data — categories the community judges important but that the data cannot yet see well. Those slots are a deliberate patch for measurement blindness. That is the taxonomy-level reading: the Top 10 is a *risk-awareness* ranking assembled from imperfect telemetry, not a measurement of your application. Treating a category's position as evidence of its true prevalence, or treating a clean scan as evidence that a category is absent, both mistake the instrument for the phenomenon.
- Does "no intrinsic oracle" mean this class of defect cannot be tested at all?No — it means the oracle must be supplied rather than derived from the observation. Once you write down, per resource type, which subject relations grant which actions, that model becomes the oracle and the category becomes assertable. The point is ordering: the modelling work has to happen first, because before it exists no amount of executing the system can decide anything. How that model is then enforced and exercised is a separate subject from how the taxonomy measures it.
- Scanners do report some access-control findings. How does that fit the argument?Those are the sub-shapes that do have an oracle: an endpoint reachable with no authentication, a directory listing, a token accepted without signature verification. No domain rule permits any of them, so the verdict is application-independent. That is exactly why the reported incidence should be read as the decidable subset plus whatever human testers contributed — a lower bound on the category rather than a measurement of it.
- Why does the OWASP Top 10 fill some slots from a community survey instead of using the contributed data for all ten?Because incidence computed from contributed testing data can only count weaknesses that current testing detects, so a category can be widespread and still under-represented. Reserving slots for practitioner judgement lets the list carry categories the community sees in real incidents before the measurement catches up. It is an acknowledged correction for the instrument's blind spot, not a popularity vote overriding evidence.
A spellchecker flags teh instantly, because no intent makes it a word. It cannot flag form where you meant from — both are valid English, and only your intent decides which is the error. Injection is the first kind of mistake; access control is the second.
saying these in an interview costs you the question
- Claiming a SAST or DAST tool finds broken access control the way it finds injection — it can only find the sub-shapes that carry their own verdict
- Reading the Top 10's incidence percentages as true prevalence rather than as detected prevalence
- Saying the category is simply "untestable" — it is undecidable without a policy model, and perfectly testable once one is supplied
- Treating the survey-sourced slots as the community outvoting the data, rather than as a deliberate correction for measurement blindness
- Assuming an authenticated 200 is evidence that the request was authorised