skip to content

How does a screen reader user move through an unfamiliar screen, and what must the screen provide for each way of moving?

level: juniorimportance: must knowfreq 74%

answer

  1. Not read from the top down
  2. Moving, not looking, builds the map
  3. Three ways: structure, control, linear
  4. A serial channel, one item at a time
  5. Each mode needs real structure

basics

~20 s

Screen reader users rarely read top to bottom. They jump between structural elements, step from one control to the next, or read linearly through everything. Each way of moving needs real structure behind the visuals, not just visual formatting.

solid answer

~50 s

A screen reader turns a screen into a serial channel — synthesised speech, or braille — delivered one item at a time, so the user builds a mental map by moving through it rather than by looking at it. Three ways of moving dominate. **By structure**: jump between section headings, major regions and groups, or pull up a generated index of every heading or every control, to reach the part worth reading. **By control**: step through the operable elements only, which is how anyone fills in a form or drives a toolbar. **Linearly**: read item after item through everything, including static text. Structural jumping only works if headings and regions are real structure rather than large bold text; stepping through controls only works if every operable thing is reachable without pointing; linear reading only works if the underlying order carries the meaning.

go deeper

for a junior

Be ready to say that the user moves through the screen rather than hearing it all, and to name the three ways of moving: by structure, by control, and linearly.

for a middle

Explain what each way of moving actually requires from the screen — real structural elements, operable things reachable without pointing, and an underlying order that matches the meaning — and give an example of each failing.

for a senior

Show that you diagnose complaints by mode: work out which way of moving broke before proposing a fix, and resist the reflex of adding description to a screen whose real problem is that it has no structure to navigate.

for a principal

Own the argument that navigability is a design property, not a remediation task: structure decided at component and layout level is what makes every screen navigable by default, and retrofitting it screen by screen never catches up.

## What a screen reader actually delivers A screen reader takes the structure an interface exposes to the platform and renders it through a **serial channel** — synthesised speech, or a refreshable braille display — one item at a time. That single fact drives everything else. The user has no peripheral view of the screen: they cannot glance at its shape, notice that there are three columns, or see that the important control sits in a corner. **Position on screen carries no information unless it is also expressed in the structure.** What the user has instead is *movement*: they build a model of the screen by travelling through it and remembering what they heard. Because the channel is serial, hearing everything is expensive. A screen with 43 items takes as long to hear as it takes to speak, and a returning user does not want to hear it twice. So experienced users do the equivalent of skimming: they move in large jumps, sample, and read in full only once they have found the part that matters. ## The three ways of moving **1. By structure.** The user jumps between the structural elements of the screen — section headings, the major regions of the layout, groups, tables, lists. Most screen readers also generate an on-demand index of every element of one type ("give me all the headings", "all the controls"), so the user picks a destination from a list instead of walking to it. This is how a returning user reaches the part they care about in two moves rather than two hundred. **2. By control.** The user steps from one operable element to the next — the same sequence a keyboard user walks, or a dedicated "next control" gesture on a touch device. This is the mode for *doing* rather than reading: filling in a form, working a toolbar, choosing from a menu. Each stop has to say what the thing is, what it is called, and whether it is currently on, off, open or closed — otherwise the user is stepping through unlabelled boxes. **3. Linearly.** The user reads item after item, or starts a continuous read, through everything on the screen including static text. This is the fallback when structure is missing, and the mode for reading something end to end. It is also the mode that exposes every ordering mistake, because it follows the underlying order exactly. | Way of moving | What the user is doing | What the screen must provide | |---|---|---| | By structure | Finding the right part of the screen | Real headings, regions and groups, correctly nested | | By control | Operating something | Every operable element reachable and self-describing | | Linearly | Reading, or coping with missing structure | An underlying order that matches the meaning | ## What each mode demands from what you build - **Structure has to be real structure, not formatting.** Text that is merely large and bold is invisible to structural jumping; the jump command goes straight past it. - **Every operable element must be reachable without pointing.** A control that responds only to a pointer never appears in the control sequence, so for this user it does not exist — even though it is plainly visible. - **Grouping matters as much as labelling.** A set of controls the eye reads as one unit is heard as unrelated items unless the grouping is in the structure. - **Decorative content must be excluded**, or linear reading fills with noise the user has to sit through. - **The order has to make sense on its own**, because the user hears the order the structure has, not the order the layout paints. ## A worked example A museum builds an audio-guide handset with 43 exhibit stops. The visual design is clean: each stop is a card with a large title, a duration and a play control. The accessibility review opens a backlog of 27 items, and the three that dominate the user complaints are all failures of a way of moving rather than missing description: 1. The stop titles are styled text rather than headings, so structural jumping has nothing to land on and the only route to stop 31 is to read past the previous 30. 2. The play control is a decorated box that reacts to a tap on its surface only, so it never appears in the sequence of controls. 3. The "back to the gallery list" action sits at the very end of the underlying order although it is painted at the top, so a linear reader meets it last. None of the three is fixed by attaching more words. They are fixed by giving the screen the structure the three modes need. That is the general lesson, and it holds identically for a handset, a desktop application and a page in a browser: **a screen reader user's experience is decided by what the screen offers to be navigated, not by how much text is attached to it.**

  • A control can only be operated by pointing at it. Which way of moving does that break, and what does the user experience?
    It breaks stepping through controls, and that is the mode people use to get things done. The control never appears in the sequence, so a user walking the controls never meets it; a user reading linearly may hear its text but has no way to operate it. The practical effect is that a visible, working feature is simply absent for that user, which is why it usually surfaces as "I cannot finish the task" rather than as a missing-label complaint.
  • Why does the generated index of every heading on a screen tell you something about the quality of that screen?
    The index is built from the structure, so it is a free read-out of what the screen actually claims to be. If it is empty, the titles are decoration. If it lists a dozen items with no shape, the sections are flat. If the entries do not describe the sections a sighted reader would name, the structure is not the one the design implies. It takes seconds and needs no product-specific knowledge to interpret.
  • If a screen has no structure at all, can a screen reader user still use it?
    Sometimes, but only by reading linearly, which is the slowest mode and scales badly. On a short, simple screen it is tolerable. On anything long or repetitive it means hearing everything each visit, with no way to skip to a section or to survey what is present. Treat linear reading as the floor a screen degrades to, never as the experience you designed for.

It is like finding a room in an unfamiliar building in the dark: you navigate by doorways, signs and landings. Take those away and you are left feeling along every wall.

saying these in an interview costs you the question

  • Says screen readers simply read a screen from top to bottom
  • Assumes the user hears the whole screen before acting
  • Thinks visual position tells the user where something is
  • Treats large bold text as a heading because it looks like one
  • Assumes a pointer-only control is fine as long as it is visible
  • Thinks adding more description compensates for missing structure