skip to content

Why does mapping a bug to CWE-20 Improper Input Validation settle nothing about the fix?

level: middleimportance: should knowfreq 44%

answer

  1. the entries carry a level
  2. Pillar, Class, Base, Variant
  3. true of nearly every flaw
  4. no resource, no consequence named
  5. the taxonomy discourages this entry

basics

~20 s

CWE-20 sits at the Class level of a four-tier taxonomy — Pillar, Class, Base, Variant. It says a check was missing, not which resource was harmed, so it constrains no change. The actionable statement lives at Base.

solid answer

~50 s

The weakness taxonomy is layered by abstraction: **Pillar** (the broadest, near-philosophical grouping), **Class** (a mistake described independent of resource and technology), **Base** (a mistake described with enough of the resource and the consequence to be testable) and **Variant** (a Base narrowed to one language, platform or technology). `CWE-20` is a Class. Almost any flaw where something arrived from outside and was then used badly can be described as improper input validation, which is exactly why the label discriminates nothing — it names an absent check without naming the resource that got hurt, and the fix is usually at the point of use rather than the point of entry. The taxonomy's own mapping guidance discourages `CWE-20` for this reason and pushes you down to a Base entry such as `CWE-787` (out-of-bounds write) or `CWE-732` (incorrect permission assignment for a critical resource), where the entry text implies a change you can verify.

go deeper

for a junior

Know that weakness entries sit at different levels of abstraction and that the broad ones are categories rather than findings. Be able to say improper input validation describes a kind of mistake, not a specific bug.

for a middle

Name the four levels in order and explain what each adds. Show why a Class-level label leaves the harmed resource and the consequence unstated, and therefore implies no change anyone can verify.

for a senior

Demonstrate the descent: work back from the consequence to a Base entry, and know both failure directions - too abstract settles nothing, too specific claims more than the evidence supports.

for a principal

Set the mapping standard the organisation files against and decide what happens to rows that cannot be mapped specifically, knowing that a register keyed on Class labels can never be counted as work.

## The four levels CWE is not a flat list of weakness names. Entries carry an **abstraction level**, and the level is the single most useful thing about an entry when someone hands you one. - **Pillar** — the broadest grouping, describing a whole family of failure in the most general terms. `CWE-664`, "Improper Control of a Resource Through its Lifetime", is a Pillar. Nothing about a Pillar is testable. It exists to organise the tree. - **Class** — a mistake stated abstractly, independent of a specific resource, language or technology. `CWE-20` (improper input validation) and `CWE-284` (improper access control) are Classes. They describe a *shape* of error. - **Base** — a mistake stated with enough of the resource and the consequence that you can decide whether a given piece of code has it. `CWE-787` (out-of-bounds write) and `CWE-732` (incorrect permission assignment for a critical resource) are Base entries. - **Variant** — a Base narrowed to a particular language, platform or technology, where that specificity changes how the mistake appears or is fixed. `CWE-121`, stack-based buffer overflow, is a Variant beneath the out-of-bounds-write Base. (The taxonomy also carries a **Compound** abstraction for weaknesses that only exist as a combination of others, and relationship concepts — chains and composites — that describe how one weakness enables another. Those are structural, not a fifth rung of specificity.) ## Why CWE-20 in particular is the trap Ask what fraction of publicly known flaws could honestly be described as "input was not validated properly". The answer is most of them. Something came from outside a trust boundary, and something downstream then did the wrong thing with it. That description is true of memory-corruption bugs, of path handling, of parsing, of deserialisation, of protocol handling — bugs whose fixes have nothing whatsoever in common. A label that is true of almost everything discriminates nothing. Concretely, `CWE-20` leaves three questions open, and they are the only three that matter for remediation: 1. **Which resource was harmed?** Memory? A file? A privilege boundary? The label does not say. 2. **What was the consequence?** A crash, a disclosure, an unintended write? The label does not say. 3. **Where does the change go?** Almost always at the point of *use*, not the point of *entry* — the sink is what defines what "valid" even means. A Class-level label pointed at the entry point actively misdirects. This is why the taxonomy's own mapping guidance marks the most abstract entries as unsuitable mapping targets: Pillars are not to be used for mapping at all, and `CWE-20` is explicitly discouraged, with a pointer to descend to something more specific. ## The two failure directions Abstraction-level errors come in two flavours and they fail differently. **Too high.** Filing at Class or Pillar produces a label everyone can agree with and nobody can act on. It also poisons counting: a hundred rows tagged `CWE-284` are a hundred unrelated changes, not one project. And it hides recurrence — the whole value of a class taxonomy is spotting that the same *specific* mistake keeps reappearing, which a bucket everything falls into cannot show. **Too low.** Filing at Variant claims more than the evidence supports. A Variant asserts a language or platform specific shape of the bug; if the reporter never saw the code and inferred the flaw from behaviour, that assertion is invented. The honest mapping is the most specific level the evidence actually supports — usually Base. ## How to descend correctly Work from the consequence back, not the input forward. Name the resource that ended up in the wrong state, name what wrongly happened to it, then find the Base entry whose text describes that. If two Base entries both fit, you probably have a chain — one weakness made a second reachable — and recording both is more honest than picking the parent Class that covers both, because the parent will tell no future reader which of the two to fix. ## The interview answer in one line `CWE-20` is a Class, and a Class names a shape of mistake rather than a resource and a consequence. It is a true statement that constrains no change, which is why it settles nothing and why the taxonomy itself tells you not to stop there.

  • Give an example of descending from CWE-20 to something actionable, and say what the descent added.
    A parser that trusts a length field and then writes past a buffer is honestly `CWE-20` and more usefully `CWE-787`, out-of-bounds write. The descent added the resource (memory) and the consequence (a write outside bounds), which together tell a reviewer what to look for and tell an engineer what change would make the finding false.
  • When is filing at Class level actually the right call?
    When the evidence genuinely does not support anything narrower — you have behaviour and no source, and two Base entries fit equally. A Class label recorded honestly, with a note that it is provisional, beats a Base label invented for tidiness. What is wrong is treating the Class label as finished work rather than as an unfinished mapping.
  • Why does a Variant exist at all if a Base already describes the mistake?
    Because for some languages and platforms the specificity changes what a reviewer looks for and what the fix is. A stack-based overflow and a heap-based one share a Base but differ in how they are found, exploited and remediated. The Variant earns its place only when that difference is real; otherwise the Base is the better mapping.

saying these in an interview costs you the question

  • Treats improper input validation as a finding rather than a category
  • Thinks all CWE entries sit at one level of specificity
  • Assumes the fix always belongs at the input boundary
  • Maps to the most specific entry available regardless of evidence
  • Believes a Pillar is a valid mapping target

context