Accessibility

An accessibility audit in one afternoon

Four hours, one engineer, no consultancy, and 212 findings in the tracker by six o'clock. Most of them were boring, which is exactly why the pass is worth running before anybody signs off a proper audit.

I have run this pass four times: twice as an employee, twice as a contractor briefed for a fortnight who finished the job in an afternoon. The most recent was a fintech product — a dashboard, an onboarding flow, a payment form, and forty other routes nobody had looked at with a keyboard since 2021. Four hours, one engineer, no consultancy. I closed the tracking board at six o'clock with 212 findings in it.

Almost none of them were interesting, which is the entire point. An afternoon does not produce a deep accessibility report. It removes the tedious, provable, mechanical failures that stand between your product and a real audit — this week rather than next quarter.

What four hours buys you

Expectations first, because the failure mode here is a report that oversells itself. One engineer with a browser, a keyboard and a scanner finds the first eighty percent of the problems, and those are overwhelmingly mechanical: a name missing, a contrast ratio of 3.1:1, a heading level skipped, a focus ring removed because it looked untidy in 2019 and nobody has touched the reset since.

You will not find the things that need assistive-technology users in the room. I cannot tell you whether the onboarding flow makes sense read aloud, or whether the validation copy is comprehensible arriving through a screen reader mid-sentence. I can tell you that the reconciliation table has no caption and no scope attributes, which is a different sentence entirely.

Triage, not an audit

"Next quarter" is ninety days of an unlabelled icon button in front of every keyboard user on the payments screen. Clearing the mechanical failures first is what stops a consultant billing an hourly rate for things a linter would have caught.

The order I do things in

The sequence matters more than the tooling: every pass changes the accessibility tree, so a later pass run too early has to be repeated. I run these six in this order.

  1. Semantics and landmarks. Exactly one main, a heading tree that reflects the visual hierarchy, and no clickable div pretending to be a control.
  2. A keyboard-only pass. Tab through every route. Where does focus go, is it always visible, does it get stuck, does it come back when the dialog closes. This finds more than any tool, needs nothing but the Tab key, and catches behaviour no scanner models.
  3. Accessible names and roles. Query every button, link, input and icon for a name that is not "button" or a bare glyph. Most turn out to be one attribute in one component.
  4. Contrast. Every text style against every background it can land on, at the token level where possible rather than instance by instance.
  5. Zoom and reflow at 400%. No horizontal scrolling for content, nothing clipped or overlapping, no fixed-height panel swallowing a button.
  6. Forms, errors and status messages. Fields labelled and associated, errors announced rather than merely coloured, async status text in a live region.

Why semantics come first

A div with a click handler becomes a button: that one change resolves the focusability, role, name and Enter-key findings, which a scanner reports as four violations and which I would otherwise write up as four tickets. It also turns the later name pass into a query over the accessibility tree rather than a hand walk of the DOM.

Where the afternoon runs out

I stopped at hour four with the reporting tables unexamined, and said so in the report rather than padding the count with guesses. I tested Chrome and Firefox on macOS plus one VoiceOver spot check, and not NVDA or JAWS on Windows — a real omission for a B2B fintech product, where much of the finance staff are on Windows laptops issued by IT. That is a hole in my data, not a claim about the product.

The 212 findings, sorted

Counts are straight from the tracker, organised by cause. Effort is per instance: what the fix took me, or what I would budget for someone who has to find the file first.

Category Findings Fix effort each Caught by automation
Missing or weak accessible names705–15 minSometimes
Contrast below 4.5:148about 5 minYes
Focus visibility and focus order3520 min – 2 hRarely
Form errors and status messages not announced221–3 hNo
Heading structure and skipped levels182 minYes
Images without a useful alt1210 min, plus content knowledgePartly
Everything else (language, skip link, ARIA misuse)7variesPartly

The distribution is the interesting part. Forty-one of the seventy name findings were the same icon-only button in the dashboard toolbar, from one component; treating that as one ticket rather than forty-one is the difference between a fix and a backlog. The scanner contributed 66 outright — the contrast and heading rows, the categories a computer can prove. The keyboard pass produced 57, including the whole focus row, and not one of those appeared in any automated report.

The form row is invisible to every scanner I ran. A scanner can see that an input has no aria-describedby; it cannot see that submitting the form with a screen reader running does nothing audible, so the user presses the button again, and again.

Automated tools: useful, not sufficient

I run axe first because it is free, fast, and genuinely excellent at a narrow class of problem. This is roughly what I did on the first afternoon, before it went into CI:

# one route per line; about 40 authenticated routes
while read -r path; do
  npx @axe-core/cli "https://app.example.com${path}" \
    --tags wcag2a,wcag2aa,wcag21aa \
    --exit
done < routes.txt

Two details in that snippet cost me time. The loop needs a session, so I passed in a cookie jar exported from a logged-in browser, because a scanner pointed at a login redirect cheerfully reports zero violations on the login page. And --exit fails the job on any violation, which is what you want in a build pipeline and not what you want on the first afternoon, when the useful output is a count.

The ratio is the thing to remember: automated checks accounted for roughly a third of the 212, and the keyboard pass found more than any tool did. Tools also generate noise. The contrast checker flagged four pieces of white text on a gradient hero it could not sample, and a third-party date picker produced three violations I could not fix without forking it. Call it an hour spent deciding that twelve results were not findings.

Writing findings people can act on

A finding that does not get fixed is a note. Every item in my report has the same five parts, in this order.

  • Route and element. /payments/new, the currency select. Reproducible in one click by someone who has never seen the code.
  • What the user experiences. "The screen reader announces 'combo box, collapsed' with no label, and the new total is not announced, so the user cannot tell whether the change took effect."
  • The smallest fix. Associate the visible label with the select, and move the total into a live region. Anything larger than a day is a discussion, not a ticket.
  • The success criterion. WCAG 2.1 SC 3.3.2 Labels or Instructions, and SC 4.1.3 Status Messages for the total.
  • Evidence. A twenty-second recording or the raw tool output, which ends the argument about reproducibility.

"The modal traps focus" is a better title than "WCAG 2.4.3 violation", and not because the standard is wrong. The engineer picking the ticket up on Thursday afternoon needs to know what to change; the product manager needs to know who is blocked. You cite the criterion anyway, because it makes the finding falsifiable and lets whoever verifies the fix check what you checked.

If a keyboard user cannot reach the button, the button does not exist, and no amount of visual polish changes that.

The header of every audit document I have written since 2022

What I would do differently next time

Three things, in the order I would change them.

  1. Start with the design tokens. Contrast is cheapest to fix centrally and most expensive per instance. My 48 findings collapsed into six token changes once I sat down with the palette — an hour I reached only after visiting fifteen pages one at a time. Opening the token file first would have bought the forms row instead.
  2. Involve one screen-reader user before writing the report. A thirty-minute session changes the shape of the document. Last time, three of my "medium" findings turned out to be noise the user already had a workflow around, and two "low" ones were the first thing they hit on every login. I had the severity roughly right and the ordering badly wrong.
  3. Put the keyboard check in CI as a smoke test. A quarterly ritual finds the same class of regression every quarter and calls it a new finding. Tab through the checkout headlessly, assert focus order and a visible outline, run axe over the five busiest routes, fail on serious violations: perhaps 150 lines. The point is not coverage, it is catching the button that quietly loses its type attribute the same day.

Four hours found 212 things, almost none of them interesting, and the product is materially better for it. What remains needs people I am not a substitute for: a Windows machine with NVDA running, and a user who does this every day. Booking that work is easier when the mechanical half of the list is already closed.

Portrait of Elliot Vance

Elliot Vance

Software engineer in Rotterdam. I help teams make their systems fast, observable and boring, usually by removing more than I add. Eleven years in, mostly on web performance, databases and platform work.

More about me