Scoring our own product in public

In three points

  • Before launch, this site went through its own scorecard, honestly, with the evidence declared.
  • The pre-launch composite is dominated by guesses, which the framework reports as low confidence, correctly.
  • The rescore, on real behavioural data, will be published as the site's first case study.

A framework that will not score its own author is a marketing asset. So before this site launched, it went through the same twelve signals, two passes, with the evidence class declared on every answer. The honest version of that exercise is worth walking through, because it demonstrates the tool's most underrated behaviour: what it does when you do not know things.

Fig. 01

Our own pre-launch scorecard, unedited

TASKACCESSBELIEFCOMMITMENT 3 measured 9 guessed, at 0.30 each COMPOSITE CONFIDENCE, PRE-LAUNCH 0.34 Most of what the author of this framework believed about his own site was opinion. The three that were measurable: the WCAG scan, keyboard traversal, and a terminology audit.
A site with no traffic has no funnel, no cohorts and no tickets, so nine of the twelve signals were judgment. The instrument said so in its own arithmetic rather than producing a flattering number, which is the behaviour it was built for.

What a pre-launch product can actually answer

A site with no traffic has no funnel telemetry, no retention cohorts, no ticket history. Of the twelve behaviour questions, exactly three could be answered with anything better than a guess: the WCAG scan (measured: the tooling runs before launch), keyboard traversal (measured: it is checked by hand, at a desk), and brand coherence (measured, if you accept a terminology audit of our own pages as counting).

Everything on the Task and Commitment surfaces (completion, recovery, navigation, activation, depth, return) is a guess until real sessions exist. The tool takes those guesses, weights them at 0.30, and returns a composite whose confidence figure is doing exactly what it was designed to do: telling the author of the framework that most of his own scorecard is currently opinion.

The part that was still useful

Two outputs survived the low confidence intact. The access and belief surfaces, where the measured answers live, produced a real reading, including one finding that was acted on before launch (contrast failures in the light theme's muted text, caught by the scan the scorecard told us to run).

And the instrument-next list came back with the Task and Commitment surfaces in full, which is the pre-launch measurement plan stated as a list: instrument scorecard starts and completions, per-question drop-off, and return visits, before trusting any number this site says about itself.

The commitment

The rescore happens when there is real behavioural data to score against, a quarter of it, so the commitment signals mean something. It gets published as a case study, with the pre-launch baseline beside it, in the same format we would use for a client: what changed, what did not work, and the delta with its confidence attached.

If the rescore comes back worse than the guesses, that gets published too. A framework whose author quietly shelves his own bad result has told you everything you need to know about the framework.

Signals this affects

All twelve, run for real. The pre-launch run leaned on contrast and legibility, assistive support and brand coherence, the three a new product can measure.

Run it on yours


Related: What a scorecard cannot tell you · Your research expired the day you shipped