Klar Barrierefrei / wcag

Transparency

How we scan and score

No black-box badge. Here is exactly what our scan does, where it stops, and how the score is derived.

Last updated: · By Paweł Dziura

The engine

We render your page in a real browser and run the pinned, versioned engine axe-core 4.12 against WCAG 2.1 and 2.2 Level AA. The exact version is recorded on every scan, so a result stays reproducible and a report stays dateable.

That the page is genuinely rendered is not a technical footnote — it is the difference between two kinds of check. Reading the source shows the HTML the server sent. Your customers see something else: what JavaScript built from it, with the menu collapsed, the product filters loaded and a cookie banner on top. That is where the barriers nobody notices actually sit. We test the state a person encounters.

Pinning the version instead of tracking the latest costs us work, and it is deliberate. A new engine version brings new rules with it, so the same shop would suddenly score differently without anything about it having changed. For comparing two scans — and for a report you may still need to explain in six months — that would be worthless.

Honest coverage: ~30–57%

Automated testing catches only part of the WCAG picture. How large a part depends on method: common industry estimates sit around 30%, while Deque’s axe-core study reaches up to 57%. We deliberately quote that range rather than a single number — anything else would be dressing it up.

The two numbers do not contradict each other — they measure different things. The industry figure near 30% counts the WCAG success criteria a machine can decide at all: roughly a third of the ruleset. Deque instead counted what share of the issues actually found in thousands of audits were caught automatically, and reaches up to 57%, because a handful of failure types repeat millions of times. Both are true. Quoting only the higher one is selling; quoting only the lower one is too.

That is why we never claim “100% compliant” from a scan. Whatever automation cannot decide with confidence, we explicitly flag for manual review — and count it visibly, rather than leaving it out of the report. A score that shows only the flattering half would be useless to an authority and to you alike.

What is checked automatically

Reliably machine-checkable, among others: colour contrast, missing alt text, form labels, link and button names, heading structure, language attributes, document titles, ARIA misuse and landmark regions.

That list looks short for “only a third of the criteria” — but it covers the bulk of the failures that actually exist. The WebAIM Million 2026 analysed one million home pages: six failure types account for 96% of everything found, and all six are on this list. The full figures are in our German guide on testing a site for accessibility. That is precisely why an automated pass is worth running even though it covers only part of the ruleset.

What needs a human

Whether an alt text is meaningful, whether focus order is logical, whether error messages are understandable, whether content is in plain language — that takes judgement. We count such items as “needs manual review” and they never lower the score.

One example makes the boundary concrete. A product image carries the alt text “image”. Mechanically everything is fine: the attribute exists, it is not empty, it is not a filename. To a screen-reader user it is worthless — she learns that an image is there, not which product she is buying. No checker decides that, because deciding it means knowing what the photo ought to show.

The same holds for keyboard operation, for video captions that exist but were auto-generated and are wrong, and for forms that are technically labelled and still incomprehensible. We name these items individually so your manual budget goes where it is needed, instead of re-walking an entire site by hand.

How the 0–100 score is computed

The score starts at 100 and subtracts a weighted penalty per confirmed issue by severity, floored at 0. Manual-review items do not lower it — we only penalise what the scan actually confirmed. A site score is the mean of its page scores.

Severity Penalty per issue
Critical −10
Serious −5
Moderate −2
Minor −1

That manual-review items never lower the score is a deliberate decision against our own commercial interest. A lower score would sell better. It would also be dishonest: at that point we do not know whether there is a problem, only that a machine cannot decide. Penalising a shop because our software reached its limit would be a statement about us, not about them.

For the same reason a site score is a plain mean of its page scores rather than a weighting by importance. Weighting would assume we know which pages matter to you — we do not. A mean is traceable and can be explained in a report; a formula with invisible factors would be exactly the black-box badge we refuse to sell.

Not an overlay

We do not modify your site or lay a widget over it. We test your real pages and tell you what to change in your own code.

Overlays promise the opposite: add one script, done. They alter the page in the browser without removing the causes in the markup — and have drawn lawsuits in several countries, because people who rely on assistive technology found the overlaid page harder to use than the original. A tool that saves you work by hiding the problem simply moves the risk onto you.

What we deliver is less convenient by design: a list of concrete findings with CSS selectors that someone has to change in your code. That takes longer than pasting in a script, but it survives scrutiny — and you own the result instead of renting it.

Frequently asked questions

Is an automated scan enough for BFSG?
No. Automated testing catches only part of WCAG issues (studies range ~30–57%). It is a strong first pass, but a manual review is needed to claim conformance — which is why we flag, never hide, what needs a human.
Which axe-core version do you run?
A pinned version (currently axe-core 4.12), recorded on every scan so a result stays reproducible and a dated report stays defensible.

Sources