skip to content

Which WCAG success criteria can an automated rule engine decide, and which need a human judgement?

level: middleimportance: must knowfreq 78%

answer

  1. Look at how the criterion is worded
  2. Some conditions are measurable, some are meaning
  3. Absence is decidable, correctness is not
  4. Silence is not the same as a pass
  5. Equivalent purpose cannot be computed

basics

~20 s

An automated rule engine decides only criteria whose pass condition is mechanically measurable — a computed value, or a required property present or absent. Criteria turning on meaning, equivalence or order need a human. That boundary belongs to the criteria.

solid answer

~50 s

Split the criteria by what their pass condition is made of. A rule engine can compute a contrast ratio, notice that an operable control exposes no accessible name, or see that a screen declares no default language — the standard states those as measurable or structural facts. It cannot decide whether a text alternative conveys the same information as the thing it replaces, whether the order controls are reached preserves the meaning of the layout, whether a heading describes the section beneath it, or whether an error message tells the user how to correct the entry. Those criteria are written in terms of **equivalence**, **purpose** and **appropriateness**, which are judgements about content. So the limit is a property of how the criteria are worded, not a maturity gap in tooling that better software will close.

go deeper

for a junior

Be ready to say that an automated accessibility check finds only some kinds of problem, and to give two examples it cannot find — a text alternative that says the wrong thing, and a heading that does not describe its section.

for a middle

An interviewer at this level expects you to explain the mechanism: point at how a success criterion is worded, and show that a measurable or structural condition is decidable while equivalence, appropriateness and preserved meaning are not.

for a senior

Show what you do with the boundary in a real delivery: which criteria you gate on every change, how needs-review output gets an owner, and how you scope the human passes by task so meaning-level criteria actually get looked at.

for a principal

Own the framing for the organisation. Decide what claim the company is entitled to make from its evidence, resist metrics that reward violation counts over usable tasks, and keep contested statistics out of the argument.

## Why the boundary belongs to the criteria, not the scanner An **automated accessibility rule engine** renders an interface, walks its structure, and evaluates a fixed set of rules against what it finds. Every rule is code, so it can settle only questions code can settle: compare a computed value against a threshold, or observe that a required property is present or absent. WCAG success criteria are not all written in that shape. Some state their pass condition as a measurable or structural fact. Others state it in terms of **meaning** — whether one thing is equivalent to another, whether an order preserves sense, whether a piece of text describes its subject. The second kind is not waiting for a better engine. It is undecidable by a rule in the same way that "is this sentence true?" is undecidable by a spelling checker: the rule has the letters, not the meaning. That is the sentence worth saying out loud in an interview — *the split is a property of how the criteria are written.* ## Reading a criterion to predict which side it falls on | How the pass condition is stated | Example success criterion | What automation can decide | |---|---|---| | A measured value against a threshold | 1.4.3 Contrast (Minimum), level AA | The criterion, wherever both values are inspectable | | A required property present or absent | 3.1.1 Language of Page, level A | The criterion | | A name that must exist **and** be right | 4.1.2 Name, Role, Value, level A | Only the missing case | | Equivalence between two contents | 1.1.1 Non-text Content, level A | Only the missing case | | An order that must preserve meaning | 2.4.3 Focus Order, level A | Nothing | | Text that must describe its subject | 2.4.6 Headings and Labels, level AA | Nothing | | A correction the system could suggest | 3.3.3 Error Suggestion, level AA | Nothing | Three families fall out of that table: 1. **Fully decidable.** The criterion is arithmetic or structural. These are the ones worth gating a build on, because they regress silently and constantly. 2. **Half decidable.** The criterion has a decidable *necessary* condition and an undecidable *sufficient* one. A rule engine finds every control with no accessible name at all; it cannot tell you the name says the wrong thing. Half a criterion reported as a pass is the most dangerous output an engine produces. 3. **Not decidable at all.** The criterion is worded with *equivalent purpose*, *preserves meaning*, *describe*, *suggestion*. No rule fires, so nothing is reported — and nothing reported reads exactly like a pass. ## The three outcomes of a run, and which one lies - **A violation.** A rule fired on something it could inspect. Strong signal, still worth confirming — a rule can fire on a construct that is fine in context. - **A needs-review item.** The engine found a construct whose correctness it cannot settle and says so. This is a work queue with an owner, not a pass. Teams routinely close these in bulk, which converts a truthful "I don't know" into a false green. - **Silence.** The largest category by far. Silence covers both "a rule checked this and it held" and "no rule covers this criterion at all", and most reports do not distinguish the two. Reading silence as coverage is the single most common mistake on this subject. ## A worked case A maintenance console for farm machinery carries a 6-step application form for a service contract: 6 steps, 23 controls, and one legacy screen at step 4 that nobody is willing to touch. An automated scan reports 0 violations and 11 needs-review items. A 40-minute human pass over the same 6 steps finds four defects: - the progress indicator at step 3 announces "Step 3" and never says how many steps there are — present, but not descriptive; - the text alternative on the machine diagram reads "diagram", which exists but is not equivalent to what the diagram shows; - at step 4 the confirm control is reached before the two fields it confirms, so the order contradicts the meaning of the screen; - the error text at step 5 reads "invalid" without saying what a valid entry looks like. Every one of those is a criterion the engine physically cannot decide, and the clean report was *accurate* about the rules it actually ran. The report was not wrong; the reading of it was. ## What to do with the boundary - Automate the fully decidable criteria and run them on **every change** — frequency is what automation is for, not depth. - Give needs-review output an owner and a recorded decision, so "I don't know" never silently becomes "fine". - Scope human passes by **task**, not by screen. Meaning-level criteria are about whether somebody can finish something. - Do not quote a percentage for how much automation catches. Say that automated rules cover part of the criteria and the rest require human judgement; the widely repeated figures have a real basis but a contested value. - When someone asks whether the product is accessible, answer in terms of **which criteria have been covered and how**, never in terms of a tool's output.

  • A rule engine reports a control has an accessible name. Which part of that success criterion is still open?
    The correctness of the name. The criterion requires user interface components to expose a name, a role and a value, and presence is the decidable half. Whether the announced name matches the visible label, describes the control's purpose, and stays right after the control changes state is judgement about content, so it needs a human pass over the actual task.
  • Your engine reports a batch of needs-review items. How should the team treat them?
    As a queue of open questions with a named owner and a recorded decision each, not as noise to suppress. A needs-review item means the rule found a construct it cannot settle — that is the engine being honest. Closing them in bulk turns an accurate 'I don't know' into a false green, and it is how meaning-level defects reach production under a clean report.
  • If tooling keeps improving, will this boundary shrink?
    The decidable side can widen at the margins as engines inspect more of what is rendered, but the criteria worded around equivalence, appropriateness and preserved meaning stay out of reach of a rule, because deciding them means understanding the content. Expect better coverage of measurable criteria and no change at all to the ones stated as judgements.

A spelling checker can prove every word exists in a dictionary and can never tell you the paragraph says the wrong thing; the limit is in the question, not in the checker.

saying these in an interview costs you the question

  • Claims a clean automated scan means the screen conforms
  • Expects better tooling to eventually automate every success criterion
  • Treats a needs-review result as a pass because nothing failed
  • Thinks a text alternative is checked by confirming one exists
  • Quotes a precise percentage for how much automation catches