What a scorecard cannot tell you
In three points
- A self-assessment is bounded by what you know: it structures your knowledge, it cannot audit it.
- It cannot produce a real tail percentile, a benchmark, or a compliance result — and ours says so on the result page.
- Its most valuable output for most teams is the list of what they could not answer.
We built a free scorecard and we would like you to run it. This article is the fine print, written large, because a measurement tool that overstates its own reach is a counterexample to everything this framework argues.
It reports what you tell it
The scorecard is a self-assessment. If your completion rate is actually 45% and you enter “60 to 80%” in good faith, the result is a well-structured presentation of a wrong belief. The tool cannot see your product; it can only organise your knowledge of it, attach honest weights to how you know each piece, and show you where the pieces disagree.
That is not a small thing — the disagreement pattern between your behaviour answers and your perception answers is robust even when the absolute numbers are shaky. But it is a different thing from measurement, and the result page prints its confidence figure for exactly this reason.
Four things it will never produce
- A real tail percentile. One respondent cannot generate a P10. The tool returns a tail-risk flag derived from your access answers, and says so. Any tool that hands a single respondent a distribution has fabricated it.
- A benchmark. There is no “industry average” on the result page, because their tail is not your tail, their weights are not your weights, and cross-company comparisons of self-reported scores compound two layers of noise into one confident-looking number.
- A compliance result. A green access surface is not WCAG conformance. Conformance is an audit against a published standard; the scorecard asks whether you have run one.
- A verdict on your strategy. A product nobody wants can score beautifully. The scorecard measures the cost of an experience and the accuracy of your beliefs about it. Whether the experience is worth building is a question it cannot see.
The output that matters most
For most teams running it the first time, the valuable output is not the composite. It is the instrument-next list — the inventory of questions you had to answer with “guessed” or “I don't have this number”. That list is your measurement roadmap, priced and ordered, and it is the part of the result that is true regardless of how optimistic your other answers were. You cannot flatter your way out of not knowing.
Run it honestly, keep the result URL, fix something, and run it again. The delta between two honest runs is worth more than the level of either.
Signals this affects
All twelve — this article is the boundary condition on the whole scoring model.
Related: Why expert review is the weakest evidence in UX · Scoring our own product in public