Scoring our own product in public

In three points

  • Before launch, this site went through its own scorecard, honestly, with the evidence declared.
  • The pre-launch composite is dominated by guesses — which the framework reports as low confidence, correctly.
  • The rescore, on real behavioural data, will be published as the site's first case study.

A framework that will not score its own author is a marketing asset. So before this site launched, it went through the same twelve signals, two passes, with the evidence class declared on every answer — and the honest version of that exercise is worth walking through, because it demonstrates the tool's most underrated behaviour: what it does when you do not know things.

What a pre-launch product can actually answer

A site with no traffic has no funnel telemetry, no retention cohorts, no ticket history. Of the twelve behaviour questions, exactly three could be answered with anything better than a guess: the WCAG scan (measured — the tooling runs before launch), keyboard traversal (measured — it is checked by hand, at a desk), and brand coherence (measured, if you accept a terminology audit of our own pages as counting).

Everything on the Task and Commitment surfaces — completion, recovery, navigation, activation, depth, return — is a guess until real sessions exist. The tool takes those guesses, weights them at 0.30, and returns a composite whose confidence figure is doing exactly what it was designed to do: telling the author of the framework that most of his own scorecard is currently opinion.

The part that was still useful

Two outputs survived the low confidence intact. The access and belief surfaces, where the measured answers live, produced a real reading — including one finding that was acted on before launch (contrast failures in the light theme's muted text, caught by the scan the scorecard told us to run).

And the instrument-next list came back with the Task and Commitment surfaces in full, which is the pre-launch measurement plan stated as a list: instrument scorecard starts and completions, per-question drop-off, and return visits, before trusting any number this site says about itself.

The commitment

The rescore happens when there is real behavioural data to score against — a quarter of it, so the commitment signals mean something. It gets published as a case study, with the pre-launch baseline beside it, in the same format we would use for a client: what changed, what did not work, and the delta with its confidence attached.

If the rescore comes back worse than the guesses, that gets published too. A framework whose author quietly shelves his own bad result has told you everything you need to know about the framework.

Signals this affects

All twelve, run for real. The pre-launch run leaned on contrast and legibility, assistive support and brand coherence — the three a new product can measure.

Run it on yours


Related: What a scorecard cannot tell you · Your research expired the day you shipped