Measured on the real web, not on fixtures.
From 2026-09-14 to 2026-09-15 we put the scanner through a controlled audit: 56 real websites, 5 pages each, scanned through production before and after a round of accuracy fixes. Then we went back to every finding the fixes removed and checked it on the live page. This is what we found, including the parts that did not flatter us.
Findings
873
from 890 · −2%
Duplicate rows
0
from 99 · the same defect, reported once
Findings with no rule id
0
from 0 · every row is filterable
September 2026 field audit II · 47 of 56 sites comparable (2 blocked the scanner in both runs, 7 changed reachability between runs) · builds d79703c → ee7f57b · ran 2026-09-14 to 2026-09-15
Engine by engine
Rule-class counts are a poor scoreboard for an accuracy fix, because a fix that gives a row an id it previously lacked moves that row between classes without removing anything. Engine totals cannot be moved by relabelling.
| Source | Before | After | Change |
|---|---|---|---|
| arc-style | 109 | 110 | +1 |
| axe-core | 202 | 213 | +11 |
| ibm-equal-access | 289 | 204 | −85 |
| lighthouse | 55 | 21 | −34 |
| pa11y | 118 | 111 | −7 |
| In-house analyzers, reported through the same field | |||
| focus-graph | 0 | 1 | +1 |
| form-behavior | 1 | 1 | 0 |
| keyboard-nav-tester | 5 | 2 | −3 |
| mobile-accessibility | 111 | 210 | +99 |
Did the fixes remove anything real?
Fewer findings proves nothing on its own. Every occurrence present in the first run and absent from the second - 1,192 of them - was re-measured on the live page with a method independent of the scanner: a verifier that shares no code with the scanner. Names, roles and hidden-ness come from Chromium's own accessibility tree over the DevTools protocol; contrast from rendered pixels, the glyphs isolated by screenshotting the text and then the same box with the text made transparent; id references by token lookup in the element's own root; focusability by calling focus() and reading the active element. Each occurrence keeps the scanner's page preparation (1920x1080, light theme, reduced motion, the scroll primer). otto.de and ynet.co.il, whose second run failed under the new coverage rule, were scanned once more alone; both failed again (renderers dying under the engines) and are counted as changed reachability, not compared. The bias runs toward reporting a lost true positive, not toward excusing the removal.
| Correctly removed | 91 |
| Still counted, just not listed | 27 |
| Same element, listed under another row | 226 |
| Re-bucketed under another rule | 2 |
| Genuinely lost | 25 |
| Removed by a documented policy | 11 |
| Unexplained, low stakes | 48 |
| Indeterminate (dynamic pages) | 762 |
| Justified - the verifier confirms the failure on the live page | 286 |
| Borderline - at exactly the 24px threshold | 0 |
| Manufactured by the fixes | 37 |
| Indeterminate (dynamic pages) | 534 |
The 25 genuinely lost occurrences are on eight sites; 16 of them are one brooklinen.com page (a rewards page whose aria-hidden panel holds focusable controls, reported by IBM's hidden-but-tabbable rules in the first run and by nothing in the second). Each is filed as an open defect against the after-build, not fixed before this page was published. The 48 unexplained rows are selectors that no longer resolve hours later on rakuten.co.jp, daum.net and stripe.com, plus two pages (linear.app/customers, n26.com's affiliate page) that hung the verifier's resolver and were skipped; they are not evidence either way. The indeterminate rows are pages the other run did not analyse, dynamic content whose selectors no longer resolve, and rows inside consent or ad vendors' frames.
Removed by a documented policy, 11 occurrences. PR #4's D44: a card link whose accessible name is the card's title while the visible text also carries the date and a 'News release' kicker (who.int). WCAG 2.5.3's Understanding document treats the heading-like text as the label, voice-control users say the title, and axe marks label-content-name-mismatch experimental; the scanner now reads the most prominent visible text run as the label. The verifier keeps the literal reading, that all visible text must be in the name, so these rows are neither counted as correctly removed nor as lost.
Where the removals came from
Only the classes the fixes touched. Everything else moved with the web, not with the scanner, and is not attributed.
| Rule class | Before | After | What changed |
|---|---|---|---|
| ibm-text_contrast_sufficient | 51 | 25 | cross-engine contrast twins fold; the painted ratio decides |
| color-contrast | 28 | 16 | pixel-verified against the page; twin rows fold |
| ibm-svg_graphics_labelled | 23 | 15 | decorative SVGs; an image's alt names the link |
| ibm-img_alt_valid | 8 | 2 | cross-engine twin rows fold |
| ibm-aria_role_valid | 7 | 2 | ARIA 1.2 roles; combobox pattern |
| skip-link-broken | 3 | 0 | a working link to an empty marker passes |
| ibm-aria_main_label_unique | 2 | 0 | cross-engine twin rows fold |
| ibm-target_spacing_sufficient | 7 | 5 | WCAG 2.5.8's 24 px circle, not the size rule |
| link-in-text-block | 11 | 9 | box-shadow underlines; GOV.UK and NHS focus styles |
| label | 4 | 2 | consent dialogs dismissed before the audit |
| label-content-name-mismatch | 7 | 6 | card links named by their title, icon-only controls |
| ibm-label_name_visible | 14 | 13 | label-in-name probe: cards, icons, form controls |
| skip-link-missing | 22 | 21 | medium when a main landmark exists |
| aria-role-invalid | 3 | 2 | ARIA 1.2 structural, graphics and DPUB roles are valid |
| landmark-main-missing | 7 | 6 | cross-engine twin rows fold |
| ibm-aria_complementary_labelled | 2 | 1 | only landmarks Chromium exposes count |
| ibm-input_checkboxes_grouped | 4 | 3 | consent dialogs dismissed before the audit |
| ibm-fieldset_label_valid | 5 | 4 | consent dialogs dismissed before the audit |
| ibm-table_headers_exists | 4 | 3 | layout tables judged by their role in the accessibility tree |
| ibm-table_headers_related | 4 | 3 | layout tables judged by their role in the accessibility tree |
| keyboard-trap | 2 | 2 | a carousel that lands focus after its slide is not a trap |
| landmark-one-main | 8 | 8 | cross-engine twin rows fold |
| landmark-no-duplicate-main | 2 | 2 | cross-engine twin rows fold |
| frame-title | 3 | 3 | rows inside ad frames are the vendor's |
| ibm-aria_hidden_nontabbable | 10 | 10 | consent traps dismissed; verifier reads the live page |
| svg-accessible-name | 8 | 9 | an image's alt names the link |
| image-alt | 7 | 8 | cross-engine twin rows fold |
| ibm-img_alt_redundant | 3 | 4 | graded low: the name is read twice, no criterion fails |
| ibm-frame_title_exists | 2 | 3 | rows inside ad frames are the vendor's |
| mobile-target-unreachable-advisory | 36 | 58 | capped at 25 per rule |
| mobile-target-size-spacing-exempt-advisory | 28 | 54 | capped at 25 per rule |
| mobile-target-size-advisory | 41 | 84 | capped at 25 per rule |
Every site
Scores moved on 38 of 47 sites: 21 up, 17 down. A score can fall because the report got more accurate - when two engines' reports of one defect merge into a single row, the row keeps the higher severity.
| Site | Findings before | After | Score |
|---|---|---|---|
| brooklinen.com | 40 | 21 | 72 → 73 |
| daum.net | 42 | 29 | 54 → 56 |
| stripe.com | 36 | 24 | 54 → 63 |
| telekom.de | 41 | 30 | 66 → 49 |
| lit.dev | 28 | 18 | 50 → 63 |
| getbootstrap.com | 36 | 29 | 52 → 54 |
| linear.app | 34 | 27 | 72 → 68 |
| coolblue.nl | 27 | 21 | 51 → 51 |
| skynewsarabia.com | 16 | 12 | 77 → 73 |
| stanford.edu | 13 | 9 | 79 → 89 |
| webflow.com | 24 | 20 | 54 → 55 |
| cam.ac.uk | 17 | 14 | 62 → 65 |
| news.ycombinator.com | 18 | 15 | 51 → 51 |
| who.int | 11 | 8 | 80 → 87 |
| bbc.com | 8 | 6 | 51 → 51 |
| framer.com | 35 | 33 | 49 → 48 |
| ghost.org | 23 | 21 | 53 → 57 |
| docs.python.org | 23 | 21 | 71 → 70 |
| berkshirehathaway.com | 10 | 9 | 86 → 86 |
| gov.sg | 11 | 10 | 85 → 85 |
| rakuten.co.jp | 36 | 35 | 49 → 48 |
| userway.org | 12 | 11 | 72 → 83 |
| designsystem.digital.gov | 8 | 7 | 79 → 89 |
| gov.pl | 13 | 13 | 73 → 70 |
| nhs.uk | 3 | 3 | 100 → 100 |
| openstreetmap.org | 19 | 19 | 50 → 59 |
| carbondesignsystem.com | 17 | 18 | 79 → 79 |
| corriere.it | 30 | 31 | 49 → 50 |
| government.nl | 2 | 3 | 89 → 89 |
| harvard.edu | 6 | 7 | 79 → 79 |
| n26.com | 10 | 11 | 58 → 63 |
| service-manual.nhs.uk | 3 | 4 | 89 → 100 |
| wise.com | 24 | 25 | 71 → 66 |
| hespress.com | 19 | 22 | 63 → 65 |
| theverge.com | 27 | 30 | 53 → 65 |
| australia.gov.au | 10 | 14 | 69 → 71 |
| about.gitlab.com | 24 | 28 | 63 → 64 |
| sanook.com | 32 | 36 | 51 → 46 |
| usa.gov | 4 | 8 | 100 → 89 |
| wix.com | 16 | 21 | 61 → 63 |
| kubernetes.io | 13 | 19 | 75 → 68 |
| ryanair.com | 15 | 21 | 64 → 56 |
| meta.discourse.org | 20 | 29 | 53 → 52 |
| djangoproject.com | 4 | 13 | 76 → 63 |
| u-tokyo.ac.jp | 15 | 24 | 83 → 73 |
| hubspot.com | 5 | 19 | 89 → 79 |
| u.ae | 10 | 25 | 77 → 64 |
What this page does not claim
- Not "2% fewer false positives." Findings fell by that much. Not every removed finding was false, and the census above is the only honest breakdown we have.
- Not a false-positive rate. Producing one requires re-verifying every finding that remains, on the live page. We re-verified what changed. A single percentage would be more quotable and less true.
- Not deterministic. Live sites ship new markup, rotate carousels and block scanners between runs. 2 sites were unreachable in both runs and are excluded; a site that changed reachability between runs is excluded too.
- Not the whole story. 45 defects were found and fixed in this audit, each with a regression test proven to fail without it. Some were in the fixes themselves, found by the census. The open items are listed in the repository.
Reproduce it
- Every number on this page is computed from the audit's raw result
files by
scripts/build-proof.pyand committed with the site; none is typed by hand. - The fixture benchmark is the other half of the evidence: precision and recall on hand-annotated ground truth.
- The scanner audited here runs on every pull request through the accessibility GitHub Action.
- 4,304 backend tests pass in both CI modes on the build that shipped these fixes.
- The raw figures for this edition: 2026-09-15.json. Other editions are listed in the field audit series.