Measured on the real web, not on fixtures.
From 2026-09-07 to 2026-09-08 we put the scanner through a controlled audit: 51 real websites, 5 pages each, scanned through production before and after a round of accuracy fixes. Then we went back to every finding the fixes removed and checked it on the live page. This is what we found, including the parts that did not flatter us.
Findings
800
from 1,158 · −31%
Duplicate rows
0
from 246 · the same defect, reported once
Findings with no rule id
0
from 128 · every row is filterable
September 2026 field audit · 45 of 51 sites comparable (5 blocked the scanner in both runs, 1 changed reachability between runs) · builds e400fec → edc3638 · ran 2026-09-07 to 2026-09-08
Every engine, down or flat
Rule-class counts are a poor scoreboard for an accuracy fix, because a fix that gives a row an id it previously lacked moves that row between classes without removing anything. Engine totals cannot be moved by relabelling. No engine's count rose.
| Source | Before | After | Change |
|---|---|---|---|
| arc-style | 74 | 61 | −13 |
| axe-core | 178 | 167 | −11 |
| ibm-equal-access | 317 | 250 | −67 |
| lighthouse | 36 | 35 | −1 |
| pa11y | 134 | 126 | −8 |
| In-house analyzers, reported through the same field | |||
| keyboard-nav-tester | 3 | 2 | −1 |
| mobile-accessibility | 416 | 159 | −257 |
Did the fixes remove anything real?
Fewer findings proves nothing on its own. Every occurrence present in
the first run and absent from the second - 703 of
them - was re-measured on the live page with a method independent of
the scanner: a grid of elementFromPoint calls, which is what a fingertip does, with no credit ever given to
an ancestor. The bias runs toward reporting a lost true positive, not
toward excusing the removal.
| Correctly removed | 337 |
| Still counted, just not listed | 35 |
| Merged into a cross-engine row | 13 |
| Re-bucketed (exempt → failing) | 32 |
| Genuinely lost | 3 |
| Unexplained, low stakes | 4 |
| Indeterminate (dynamic pages) | 279 |
| Justified - undersized and crowded | 23 |
| Borderline - at exactly the 24px threshold | 3 |
| Manufactured by the fixes | 1 |
| Indeterminate (dynamic pages) | 44 |
The 3 genuinely lost findings became a defect of their own, fixed and deployed before this page was published. The indeterminate rows are carousels and news feeds whose selectors no longer resolve hours later; they are not evidence either way, and we do not count them as either.
Where the removals came from
Only the classes the fixes touched. Everything else moved with the web, not with the scanner, and is not attributed.
| Rule class | Before | After | What changed |
|---|---|---|---|
| mobile-target-size-advisory | 213 | 73 | hit area counts overlays and descendants |
| (no rule id) | 128 | 0 | rule_id stamped on every row |
| mobile-target-unreachable-advisory | 104 | 35 | reveal-on-focus, scrollable ancestors |
| mobile-target-size-spacing-exempt-advisory | 83 | 44 | hit area counts overlays and descendants |
| ibm-text_contrast_sufficient | 63 | 45 | cross-engine contrast dedup |
| ibm-label_name_visible | 18 | 8 | label-in-name probe wired |
| mobile-target-size-aa-fail | 16 | 7 | hit area counts overlays and descendants |
| ibm-a_text_purpose | 10 | 1 | link-name canonical dedup |
| ibm-aria_complementary_labelled | 15 | 6 | lone-landmark probe |
| ibm-label_ref_valid | 10 | 1 | IDREF ground-truth probe |
| skip-link-missing | 15 | 7 | skip-link detection widened |
| ibm-svg_graphics_labelled | 28 | 23 | decorative-SVG probe |
| ibm-aria_role_valid | 9 | 6 | ARIA 1.2 combobox |
| ibm-aria_id_unique | 25 | 22 | IDREF ground-truth probe |
Every site
Scores moved on 25 of 45 sites: 19 up, 6 down. A score can fall because the report got more accurate - when two engines' reports of one defect merge into a single row, the row keeps the higher severity.
| Site | Findings before | After | Score |
|---|---|---|---|
| astro.build | 30 | 7 | 71 → 88 |
| theguardian.com | 44 | 24 | 51 → 51 |
| spiegel.de | 51 | 31 | 54 → 60 |
| ikea.com | 35 | 17 | 75 → 75 |
| naver.com | 41 | 24 | 51 → 54 |
| aljazeera.net | 32 | 16 | 75 → 61 |
| allbirds.com | 32 | 16 | 56 → 59 |
| gymshark.com | 54 | 38 | 51 → 51 |
| nextjs.org | 54 | 38 | 50 → 66 |
| nuxt.com | 51 | 38 | 52 → 52 |
| service-public.fr | 18 | 6 | 89 → 89 |
| svelte.dev | 25 | 13 | 79 → 79 |
| angular.dev | 38 | 27 | 63 → 67 |
| ourworldindata.org | 47 | 36 | 54 → 54 |
| mit.edu | 27 | 17 | 76 → 79 |
| react.dev | 37 | 27 | 55 → 60 |
| remix.run | 26 | 16 | 61 → 74 |
| d3js.org | 33 | 24 | 56 → 56 |
| docusaurus.io | 25 | 16 | 71 → 75 |
| unipd.it | 44 | 35 | 57 → 57 |
| lemonde.fr | 33 | 25 | 75 → 75 |
| bbc.com | 23 | 16 | 78 → 74 |
| developer.mozilla.org | 20 | 13 | 82 → 82 |
| accessinfocus.com | 16 | 10 | 85 → 84 |
| design-system.service.gov.uk | 10 | 4 | 89 → 89 |
| squidfunk.github.io | 34 | 28 | 67 → 71 |
| canada.ca | 14 | 9 | 83 → 87 |
| gov.uk | 10 | 5 | 87 → 89 |
| sarasoueidan.com | 14 | 9 | 75 → 79 |
| webaim.org | 10 | 5 | 86 → 87 |
| a11yproject.com | 14 | 10 | 83 → 86 |
| about.readthedocs.com | 17 | 13 | 50 → 52 |
| techcrunch.com | 52 | 48 | 52 → 51 |
| horizonreaches.com | 11 | 8 | 79 → 79 |
| allaccessworld.store | 5 | 3 | 83 → 89 |
| asahi.com | 32 | 30 | 51 → 47 |
| drupal.org | 8 | 6 | 79 → 79 |
| alphagov.github.io | 40 | 38 | 57 → 57 |
| tetralogical.com | 5 | 3 | 89 → 89 |
| wordpress.org | 8 | 6 | 89 → 89 |
| bund.de | 5 | 4 | 89 → 89 |
| inclusive-components.design | 8 | 7 | 85 → 85 |
| vuejs.org | 17 | 16 | 57 → 61 |
| accessibilitypro.app | 1 | 1 | 100 → 100 |
| european-union.europa.eu | 7 | 17 | 89 → 75 |
What this page does not claim
- Not "31% fewer false positives." Findings fell by that much. Not every removed finding was false, and the census above is the only honest breakdown we have.
- Not a false-positive rate. Producing one requires re-verifying every finding that remains, on the live page. We re-verified what changed. A single percentage would be more quotable and less true.
- Not deterministic. Live sites ship new markup, rotate carousels and block scanners between runs. 5 sites were unreachable in both runs and are excluded; a site that changed reachability between runs is excluded too.
- Not the whole story. 40 defects were found and fixed in this audit, each with a regression test proven to fail without it. Some were in the fixes themselves, found by the census. The open items are listed in the repository.
Reproduce it
- Every number on this page is computed from the audit's raw result
files by
scripts/build-proof.pyand committed with the site; none is typed by hand. - The fixture benchmark is the other half of the evidence: precision and recall on hand-annotated ground truth.
- The scanner audited here runs on every pull request through the accessibility GitHub Action.
- 3,684 backend tests pass in both CI modes on the build that shipped these fixes.
- The raw figures for this edition: 2026-09.json. Other editions are listed in the field audit series.