skip to content

The OWASP Top 10 is often treated as a checklist to test an application against. Explain what the list actually is, how its entries are derived and ranked, and what a team is getting wrong when it says "we tested for the Top 10, so we are secure".

level: middleimportance: must knowfreq 68%

answer

  1. categories, not bugs
  2. incidence rate per app, not finding counts
  3. survey slots = risks the data can't see
  4. population statistic ≠ your risk order
  5. awareness doc; ASVS is the verification standard

basics

~20 s

An awareness document: ten broad risk categories, each rolling up many underlying weakness types, ranked mainly by how often they appear in tested applications weighted by exploitability and impact. It is shared vocabulary and a coverage prompt, not a verification standard or a sufficient requirements list.

solid answer

~60 s

It is a **risk-category taxonomy**, not a bug list. Each entry — broken access control, cryptographic failures, injection, insecure design, security misconfiguration, vulnerable and outdated components, identification and authentication failures, software and data integrity failures, logging and monitoring failures, server-side request forgery — aggregates dozens of underlying weakness types. The ranking is a **population statistic**: most slots come from measured incidence across a large corpus of tested applications, weighted by exploitability and impact; a small number of slots come from a practitioner survey precisely because some risks (design flaws, SSRF) are systematically under-reported by automated testing rather than genuinely rare. Two things follow. The ranking describes the industry, not your application — your worst risk may be ninth on the list or absent from it. And you cannot test a *category*; you verify *controls*. "Tested for the Top 10" usually means "ran a scanner", which is strongest exactly where the taxonomy is weakest. Correct uses: training curriculum, triage vocabulary, and a completeness check over a threat model. Verification belongs to a requirements standard such as ASVS.

go deeper

for a junior

Name it as an awareness list of the ten most common web application risk categories, list several, and say it is a starting point rather than a complete requirements set.

for a middle

Explain the derivation: incidence rate across tested applications plus exploitability and impact weighting, with a couple of survey-driven slots for risks the data under-reports. Stress category-versus-bug.

for a senior

Focus on misuse: ranking is a population statistic, categories are not testable units, and the tooling that certifies 'we tested the Top 10' is biased toward the categories the taxonomy handles worst. Point to a verification standard for requirements.

for a principal

Frame it as a governance artefact: it gives an organisation a shared risk vocabulary and a cheap completeness check on threat models, while requirements, verification depth and remediation order must come from the organisation's own risk model.

## What the list is The OWASP Top 10 is an awareness document published by a vendor-neutral non-profit. It names ten *categories* of application security risk, in rank order. A category is not a vulnerability: "broken access control" covers everything from an unchecked object identifier in a URL to a missing tenant filter in a background job to a privileged endpoint reachable without a role check. Underneath each category sit many catalogued weakness types (the CWE catalogue is the usual mapping target), and one category can span more than thirty of them. ## How entries are derived Most slots are **data-driven**. Contributors submit findings from large numbers of tested applications; for each weakness type the compilers compute an *incidence rate* — the share of applications in which that weakness was found at least once — rather than raw finding counts, so an app with 400 instances of one bug does not outvote 400 apps with one instance each. That rate is combined with weightings for exploitability and technical impact drawn from the weakness catalogue. A minority of slots come from a **community survey**. This is the most misunderstood part of the methodology and the most important. Automated and even manual testing can only report what it can *see*; a control that was never designed leaves no signature to scan for, and a server-side request-forgery path may never be probed by a test harness. Survey slots exist to admit risks the data is structurally blind to. That is why "insecure design" and SSRF entered as categories: not because they are newly invented, but because the measurement method under-counts them. ## Why ranking is not your ranking The position of a category reflects the *sampled population* of applications, which skews toward organisations that pay for testing and toward whatever technology mix dominated the collection window. Your application's risk is the product of your architecture, your data sensitivity and your exposure. A payments back office with no browser surface has a different profile from a public content site. Treating the numbering as a remediation order is the most common misuse after checklist thinking. ## Why "we tested for the Top 10" is a category error Three reasons. First, **a category is not testable**. You test a control: "every object-scoped endpoint enforces ownership server-side". Ten categories yield hundreds of concrete controls, and the mapping from one to the other is the work. Second, **coverage is a floor, not a ceiling**. Ten items were chosen because they are common, and commonality is precisely what makes them a poor proxy for *your* worst case — business-logic abuse, tenancy isolation and domain-specific fraud paths rarely appear as taxonomy entries. Third, **the tooling that produces the claim is biased in the same direction as the data**. Scanners find misconfiguration, known-vulnerable components and reflected injection well; they find broken access control and insecure design poorly, because both require knowing the intended policy, which is not in the code. ## What the list is good for As a **common vocabulary** it lets a developer, a tester and a risk owner name the same defect the same way, and it makes findings comparable across teams and vendors. As a **coverage prompt** it is a cheap completeness check at the end of a threat model: walk the ten categories and ask whether the design has anything to say about each — that catches the silent gaps (nobody thought about outbound request destinations; nobody defined what gets logged on an authorisation failure). As a **curriculum** it is a reasonable ordering for teaching. What it is not is a compliance baseline; the same organisation publishes a verification standard (ASVS) with levelled, testable requirements, and a maturity model for programme design, precisely because the Top 10 was never meant to carry that load.

  • Why is incidence rate per application used instead of the total number of findings?
    Total findings are dominated by whatever a scanner emits in bulk — one application with a thousand instances of the same pattern would swamp the dataset. Incidence rate asks how many distinct applications exhibit the weakness at least once, which better estimates the probability that an arbitrary new application has it. It deliberately trades away severity-of-density information, which is why exploitability and impact weightings are applied separately.
  • If not the Top 10, what would you actually give a team as security requirements?
    A levelled verification standard such as ASVS, trimmed to the application's risk tier, plus threat-model output specific to the system. That yields testable statements — session tokens are rotated on privilege change, every object access is authorised server-side — which can be turned into tests and review criteria. The Top 10 then serves as the coverage check that the trimmed requirement set did not silently drop a whole category.
  • Which Top 10 categories are systematically under-found by automated tooling, and why?
    Broken access control and insecure design, because both require knowing the *intended* policy: a scanner sees that a request succeeded, not that it should have been denied. Software and data integrity failures are also weak spots since they involve build and update pipelines outside the running app. Conversely, vulnerable components, misconfiguration and reflected injection are the categories tools handle best, which biases any tool-derived risk picture.

saying these in an interview costs you the question

  • Calling it a standard or claiming compliance with it
  • Reading the numbering as a remediation priority order for their own app
  • Believing each entry is a single vulnerability rather than an aggregate of many weakness types
  • Assuming the ranking comes purely from survey opinion, or purely from data — it is deliberately both
  • Treating a clean scanner report against the ten categories as evidence of security

context