Measured on the real web, not on fixtures.
From 2026-09-15 to 2026-09-16 we put the scanner through a controlled audit: 56 real websites, 5 pages each, scanned through production before and after a round of accuracy fixes. Then we went back to every finding the fixes removed and checked it on the live page. This is what we found, including the parts that did not flatter us.
Findings
928
from 926 · 0%
Duplicate rows
0
from 0 · the same defect, reported once
Findings with no rule id
0
from 0 · every row is filterable
September 2026 field audit III · 50 of 56 sites comparable (2 blocked the scanner in both runs, 4 changed reachability between runs) · builds ee7f57b → 115e5b2 · ran 2026-09-15 to 2026-09-16
Engine by engine
Rule-class counts are a poor scoreboard for an accuracy fix, because a fix that gives a row an id it previously lacked moves that row between classes without removing anything. Engine totals cannot be moved by relabelling.
| Source | Before | After | Change |
|---|---|---|---|
| arc-style | 114 | 119 | +5 |
| axe-core | 225 | 237 | +12 |
| ibm-equal-access | 219 | 198 | −21 |
| lighthouse | 22 | 22 | 0 |
| pa11y | 113 | 121 | +8 |
| In-house analyzers, reported through the same field | |||
| focus-graph | 2 | 1 | −1 |
| form-behavior | 1 | 1 | 0 |
| keyboard-nav-tester | 2 | 2 | 0 |
| mobile-accessibility | 228 | 227 | −1 |
Did the fixes remove anything real?
Fewer findings proves nothing on its own. Every occurrence present in the first run and absent from the second - 1,176 of them - was re-measured on the live page with a method independent of the scanner: a verifier that shares no code with the scanner. Names, roles and hidden-ness come from Chromium's own accessibility tree over the DevTools protocol; contrast from rendered pixels, the glyphs isolated by screenshotting the text and then the same box with the text made transparent; id references by token lookup in the element's own root; focusability by calling focus() and reading the active element. Each occurrence keeps the scanner's page preparation (1920x1080, light theme, reduced motion, the scroll primer). Every occurrence the comparison counted has a row here: target-size occurrences, which the verifier hands to a separate mobile pass, are counted as indeterminate rather than dropped. rakuten.co.jp and stanford.edu lost three pages each to task timeouts while running side by side and were scanned once more alone; those runs are the ones compared. The bias runs toward reporting a lost true positive, not toward excusing the removal.
| Correctly removed | 42 |
| Still counted, just not listed | 31 |
| Same element, listed under another row | 123 |
| Re-bucketed under another rule | 5 |
| Genuinely lost | 13 |
| Removed by a documented policy | 12 |
| Indeterminate (dynamic pages)includes 19 not explained | 950 |
| Justified - the verifier confirms the failure on the live page | 232 |
| Borderline - at exactly the 24px threshold | 0 |
| Manufactured by the fixes | 42 |
| Indeterminate (dynamic pages) | 587 |
The 13 genuinely lost occurrences are on six sites: hespress.com's consent widget (four icon-only buttons whose visible Material-icon ligature the verifier reads as a label, and a tabbable span), a skip link on one gitlab.com page, an unnamed svg each on linear.app and lit.dev's playground, a 'Like post' button on framer.com whose visible count is not in its name, and a skipped heading level and a tabbable div on theverge.com. Each is filed as an open defect against the after-build, not fixed before this page was published. The 19 unexplained rows are occurrences the reconciliation could not resolve on the live page hours later (usa.gov, gitlab.com, theverge.com and four others) and are shown inside Indeterminate; they are not evidence either way. The indeterminate rows are pages the other run did not analyse, target-size occurrences (judged by a separate mobile pass this census does not fold in), dynamic content whose selectors no longer resolve, and rows inside consent or ad vendors' frames.
Removed by a documented policy, 12 occurrences. PR #12's D78: brooklinen.com opens Attentive's SMS signup about twenty seconds after load, moves focus into its iframe and sets aria-hidden on the navigation, the nano-bar and main. The second run audited the rewards page in that state and reported the page's own chrome as focusable-but-hidden; the scanner now refuses the overlay's host, so those rows are gone. The verifier does not block it, still meets the popup, and keeps the literal reading, so these rows are neither counted as correctly removed nor as lost.
Where the removals came from
Only the classes the fixes touched. Everything else moved with the web, not with the scanner, and is not attributed.
| Rule class | Before | After | What changed |
|---|---|---|---|
| ibm-text_contrast_sufficient | 28 | 24 | the same locator for IBM's XPaths |
| ibm-aria_id_unique | 4 | 1 | locations inside shadow roots resolve; a collapsed trigger's aria-controls is verified, not shipped on the engine's word |
| ibm-svg_graphics_labelled | 15 | 12 | an svg inside a named control in a shadow root is contradicted |
| aria-hidden-focus | 5 | 4 | the same overlays and ad frames |
| ibm-element_tabbable_role_valid | 29 | 28 | tiles an overlay had hidden are reported again |
| frame-title | 3 | 3 | ad frames that never load carry no untitled iframe |
| color-contrast | 17 | 19 | a location that matches several elements is measured on the one it names or not at all |
| ibm-aria_hidden_nontabbable | 10 | 13 | ad players and timed marketing overlays no longer load for the engines; rows inside them are gone, rows the overlay hid are back |
Every site
Scores moved on 21 of 50 sites: 12 up, 9 down. A score can fall because the report got more accurate - when two engines' reports of one defect merge into a single row, the row keeps the higher severity.
| Site | Findings before | After | Score |
|---|---|---|---|
| theverge.com | 30 | 23 | 65 → 58 |
| sanook.com | 36 | 31 | 46 → 48 |
| usa.gov | 8 | 4 | 89 → 100 |
| linear.app | 27 | 24 | 68 → 71 |
| hespress.com | 22 | 20 | 65 → 63 |
| n26.com | 11 | 9 | 63 → 63 |
| u-tokyo.ac.jp | 24 | 22 | 73 → 79 |
| bombas.com | 24 | 23 | 49 → 50 |
| djangoproject.com | 13 | 12 | 63 → 63 |
| about.gitlab.com | 28 | 27 | 64 → 64 |
| lit.dev | 18 | 17 | 63 → 63 |
| ryanair.com | 21 | 20 | 56 → 56 |
| userway.org | 11 | 10 | 83 → 88 |
| designsystem.digital.gov | 7 | 6 | 89 → 89 |
| australia.gov.au | 14 | 14 | 71 → 71 |
| bbc.com | 6 | 6 | 51 → 51 |
| getbootstrap.com | 29 | 29 | 54 → 60 |
| cam.ac.uk | 14 | 14 | 65 → 62 |
| carbondesignsystem.com | 18 | 18 | 79 → 79 |
| coolblue.nl | 21 | 21 | 51 → 52 |
| framer.com | 33 | 33 | 48 → 49 |
| ghost.org | 21 | 21 | 57 → 57 |
| github.com | 12 | 12 | 83 → 83 |
| gov.pl | 13 | 13 | 70 → 70 |
| gov.sg | 10 | 10 | 85 → 85 |
| government.nl | 3 | 3 | 89 → 89 |
| news.ycombinator.com | 15 | 15 | 51 → 51 |
| harvard.edu | 7 | 7 | 79 → 79 |
| kubernetes.io | 19 | 19 | 68 → 68 |
| service-manual.nhs.uk | 4 | 4 | 100 → 100 |
| nhs.uk | 3 | 3 | 100 → 100 |
| docs.python.org | 21 | 21 | 70 → 70 |
| rakuten.co.jp | 35 | 35 | 48 → 48 |
| skynewsarabia.com | 12 | 12 | 73 → 73 |
| telekom.de | 30 | 30 | 49 → 50 |
| webflow.com | 20 | 20 | 55 → 55 |
| who.int | 8 | 8 | 87 → 87 |
| wise.com | 25 | 25 | 66 → 66 |
| wix.com | 21 | 21 | 63 → 63 |
| corriere.it | 31 | 32 | 50 → 46 |
| meta.discourse.org | 29 | 30 | 52 → 53 |
| ethz.ch | 11 | 12 | 79 → 79 |
| hubspot.com | 19 | 20 | 79 → 79 |
| stanford.edu | 9 | 10 | 89 → 74 |
| en.wikipedia.org | 31 | 32 | 59 → 53 |
| stripe.com | 24 | 26 | 63 → 59 |
| berkshirehathaway.com | 9 | 12 | 86 → 82 |
| daum.net | 29 | 34 | 56 → 63 |
| openstreetmap.org | 19 | 24 | 59 → 59 |
| brooklinen.com | 21 | 34 | 73 → 61 |
What this page does not claim
- Not "0% fewer false positives." Findings fell by that much. Not every removed finding was false, and the census above is the only honest breakdown we have.
- Not a false-positive rate. Producing one requires re-verifying every finding that remains, on the live page. We re-verified what changed. A single percentage would be more quotable and less true.
- Not deterministic. Live sites ship new markup, rotate carousels and block scanners between runs. 2 sites were unreachable in both runs and are excluded; a site that changed reachability between runs is excluded too.
- Not the whole story. 6 defects were found and fixed in this audit, each with a regression test proven to fail without it. Some were in the fixes themselves, found by the census. The open items are listed in the repository.
Reproduce it
- Every number on this page is computed from the audit's raw result
files by
scripts/build-proof.pyand committed with the site; none is typed by hand. - The fixture benchmark is the other half of the evidence: precision and recall on hand-annotated ground truth.
- The scanner audited here runs on every pull request through the accessibility GitHub Action.
- 4,350 backend tests pass in both CI modes on the build that shipped these fixes.
- The raw figures for this edition: 2026-09-16.json. Other editions are listed in the field audit series.