Skip to main content
Real web · Before and after · Re‑verified live

Measured on the real web, not on fixtures.

From 2026-09-14 to 2026-09-15 we put the scanner through a controlled audit: 56 real websites, 5 pages each, scanned through production before and after a round of accuracy fixes. Then we went back to every finding the fixes removed and checked it on the live page. This is what we found, including the parts that did not flatter us.

Findings

873

from 890 · −2%

Duplicate rows

0

from 99 · the same defect, reported once

Findings with no rule id

0

from 0 · every row is filterable

September 2026 field audit II · 47 of 56 sites comparable (2 blocked the scanner in both runs, 7 changed reachability between runs) · builds d79703cee7f57b · ran 2026-09-14 to 2026-09-15

Engine by engine

Rule-class counts are a poor scoreboard for an accuracy fix, because a fix that gives a row an id it previously lacked moves that row between classes without removing anything. Engine totals cannot be moved by relabelling.

SourceBeforeAfterChange
arc-style109110+1
axe-core202213+11
ibm-equal-access289204−85
lighthouse5521−34
pa11y118111−7
In-house analyzers, reported through the same field
focus-graph01+1
form-behavior110
keyboard-nav-tester52−3
mobile-accessibility111210+99

Did the fixes remove anything real?

Fewer findings proves nothing on its own. Every occurrence present in the first run and absent from the second - 1,192 of them - was re-measured on the live page with a method independent of the scanner: a verifier that shares no code with the scanner. Names, roles and hidden-ness come from Chromium's own accessibility tree over the DevTools protocol; contrast from rendered pixels, the glyphs isolated by screenshotting the text and then the same box with the text made transparent; id references by token lookup in the element's own root; focusability by calling focus() and reading the active element. Each occurrence keeps the scanner's page preparation (1920x1080, light theme, reduced motion, the scroll primer). otto.de and ynet.co.il, whose second run failed under the new coverage rule, were scanned once more alone; both failed again (renderers dying under the engines) and are counted as changed reachability, not compared. The bias runs toward reporting a lost true positive, not toward excusing the removal.

1,192 removed occurrences
Correctly removed91
Still counted, just not listed27
Same element, listed under another row226
Re-bucketed under another rule2
Genuinely lost25
Removed by a documented policy11
Unexplained, low stakes48
Indeterminate (dynamic pages)762
857 newly reported AA failures
Justified - the verifier confirms the failure on the live page286
Borderline - at exactly the 24px threshold0
Manufactured by the fixes37
Indeterminate (dynamic pages)534

The 25 genuinely lost occurrences are on eight sites; 16 of them are one brooklinen.com page (a rewards page whose aria-hidden panel holds focusable controls, reported by IBM's hidden-but-tabbable rules in the first run and by nothing in the second). Each is filed as an open defect against the after-build, not fixed before this page was published. The 48 unexplained rows are selectors that no longer resolve hours later on rakuten.co.jp, daum.net and stripe.com, plus two pages (linear.app/customers, n26.com's affiliate page) that hung the verifier's resolver and were skipped; they are not evidence either way. The indeterminate rows are pages the other run did not analyse, dynamic content whose selectors no longer resolve, and rows inside consent or ad vendors' frames.

Removed by a documented policy, 11 occurrences. PR #4's D44: a card link whose accessible name is the card's title while the visible text also carries the date and a 'News release' kicker (who.int). WCAG 2.5.3's Understanding document treats the heading-like text as the label, voice-control users say the title, and axe marks label-content-name-mismatch experimental; the scanner now reads the most prominent visible text run as the label. The verifier keeps the literal reading, that all visible text must be in the name, so these rows are neither counted as correctly removed nor as lost.

Where the removals came from

Only the classes the fixes touched. Everything else moved with the web, not with the scanner, and is not attributed.

Rule classBeforeAfterWhat changed
ibm-text_contrast_sufficient5125cross-engine contrast twins fold; the painted ratio decides
color-contrast2816pixel-verified against the page; twin rows fold
ibm-svg_graphics_labelled2315decorative SVGs; an image's alt names the link
ibm-img_alt_valid82cross-engine twin rows fold
ibm-aria_role_valid72ARIA 1.2 roles; combobox pattern
skip-link-broken30a working link to an empty marker passes
ibm-aria_main_label_unique20cross-engine twin rows fold
ibm-target_spacing_sufficient75WCAG 2.5.8's 24 px circle, not the size rule
link-in-text-block119box-shadow underlines; GOV.UK and NHS focus styles
label42consent dialogs dismissed before the audit
label-content-name-mismatch76card links named by their title, icon-only controls
ibm-label_name_visible1413label-in-name probe: cards, icons, form controls
skip-link-missing2221medium when a main landmark exists
aria-role-invalid32ARIA 1.2 structural, graphics and DPUB roles are valid
landmark-main-missing76cross-engine twin rows fold
ibm-aria_complementary_labelled21only landmarks Chromium exposes count
ibm-input_checkboxes_grouped43consent dialogs dismissed before the audit
ibm-fieldset_label_valid54consent dialogs dismissed before the audit
ibm-table_headers_exists43layout tables judged by their role in the accessibility tree
ibm-table_headers_related43layout tables judged by their role in the accessibility tree
keyboard-trap22a carousel that lands focus after its slide is not a trap
landmark-one-main88cross-engine twin rows fold
landmark-no-duplicate-main22cross-engine twin rows fold
frame-title33rows inside ad frames are the vendor's
ibm-aria_hidden_nontabbable1010consent traps dismissed; verifier reads the live page
svg-accessible-name89an image's alt names the link
image-alt78cross-engine twin rows fold
ibm-img_alt_redundant34graded low: the name is read twice, no criterion fails
ibm-frame_title_exists23rows inside ad frames are the vendor's
mobile-target-unreachable-advisory3658capped at 25 per rule
mobile-target-size-spacing-exempt-advisory2854capped at 25 per rule
mobile-target-size-advisory4184capped at 25 per rule

Every site

Scores moved on 38 of 47 sites: 21 up, 17 down. A score can fall because the report got more accurate - when two engines' reports of one defect merge into a single row, the row keeps the higher severity.

SiteFindings beforeAfterScore
brooklinen.com402172 → 73
daum.net422954 → 56
stripe.com362454 → 63
telekom.de413066 → 49
lit.dev281850 → 63
getbootstrap.com362952 → 54
linear.app342772 → 68
coolblue.nl272151 → 51
skynewsarabia.com161277 → 73
stanford.edu13979 → 89
webflow.com242054 → 55
cam.ac.uk171462 → 65
news.ycombinator.com181551 → 51
who.int11880 → 87
bbc.com8651 → 51
framer.com353349 → 48
ghost.org232153 → 57
docs.python.org232171 → 70
berkshirehathaway.com10986 → 86
gov.sg111085 → 85
rakuten.co.jp363549 → 48
userway.org121172 → 83
designsystem.digital.gov8779 → 89
gov.pl131373 → 70
nhs.uk33100 → 100
openstreetmap.org191950 → 59
carbondesignsystem.com171879 → 79
corriere.it303149 → 50
government.nl2389 → 89
harvard.edu6779 → 79
n26.com101158 → 63
service-manual.nhs.uk3489 → 100
wise.com242571 → 66
hespress.com192263 → 65
theverge.com273053 → 65
australia.gov.au101469 → 71
about.gitlab.com242863 → 64
sanook.com323651 → 46
usa.gov48100 → 89
wix.com162161 → 63
kubernetes.io131975 → 68
ryanair.com152164 → 56
meta.discourse.org202953 → 52
djangoproject.com41376 → 63
u-tokyo.ac.jp152483 → 73
hubspot.com51989 → 79
u.ae102577 → 64

What this page does not claim

  • Not "2% fewer false positives." Findings fell by that much. Not every removed finding was false, and the census above is the only honest breakdown we have.
  • Not a false-positive rate. Producing one requires re-verifying every finding that remains, on the live page. We re-verified what changed. A single percentage would be more quotable and less true.
  • Not deterministic. Live sites ship new markup, rotate carousels and block scanners between runs. 2 sites were unreachable in both runs and are excluded; a site that changed reachability between runs is excluded too.
  • Not the whole story. 45 defects were found and fixed in this audit, each with a regression test proven to fail without it. Some were in the fixes themselves, found by the census. The open items are listed in the repository.

Reproduce it

  • Every number on this page is computed from the audit's raw result files by scripts/build-proof.py and committed with the site; none is typed by hand.
  • The fixture benchmark is the other half of the evidence: precision and recall on hand-annotated ground truth.
  • The scanner audited here runs on every pull request through the accessibility GitHub Action.
  • 4,304 backend tests pass in both CI modes on the build that shipped these fixes.
  • The raw figures for this edition: 2026-09-15.json. Other editions are listed in the field audit series.