Why expert review is the weakest evidence in UX, and why we weight it 0.30
In three points
- RUCF weights evidence numerically: observed 1.00, reported 0.70, inferred 0.50, assumed 0.30.
- Expert judgment — including the author's — sits at the bottom, and that is the correct place for it.
- AI classification enters at 0.50 and stays there until a human validates a sample.
Medicine solved this problem decades ago. The GRADE system ranks evidence by how it was gathered: randomised trials above observational studies, observational studies above expert opinion. Nobody in medicine finds it insulting that a professor's opinion ranks below a trial. It is simply what the opinion is worth as evidence.
UX has the same hierarchy and refuses to write it down. Everyone in the field knows that a heuristic review is weaker than a moderated test, which is weaker than production telemetry. We say it in hallways and omit it from frameworks — because the frameworks are sold by the people doing the reviews.
The RUCF weights
RUCF makes the hierarchy arithmetic. Observed behaviour — telemetry, session data, moderated task tests — enters at 1.00. Reported experience — surveys, interviews, tickets — at 0.70, because it passes through memory and self-presentation on the way to you. Anything inferred — analytics proxies, AI classification of tickets or replays — at 0.50. Expert judgment, including a senior designer's heuristic review, including mine, at 0.30.
That 0.30 is not a slight. It is the observed reliability of expert prediction in a domain with fast feedback distortion: experts predict where users will struggle at rates that are useful for generating hypotheses and hopeless for confirming them. A review is a list of places to point instruments. It is not the instrument.
The AI clause
The 0.50 tier deserves its own paragraph, because it is where the industry is currently misfiling the most confidence. A model that classifies ten thousand support tickets by theme is genuinely useful — volume is precisely what models are good at. The failure mode is nuance, and it is confident rather than obvious: the misclassifications do not look like errors, they look like findings.
So in RUCF, anything a model classified enters at 0.50 and stays there until a human validates a sample — at which point it is promoted to the class of the validation method, not to 1.00 by default. A framework that discounts AI evidence in its own arithmetic is more trustworthy than one that markets AI as a feature.
What the weight changes in practice
Confidence multiplies into priority: recoverable cost × confidence ÷ effort. A large cost supported only by expert opinion ranks below a smaller cost you observed, which is the opposite of how most roadmaps are argued. The weight is what stops a persuasive review from outranking an unglamorous funnel export — and it applies to the framework's author as firmly as to anyone else.
Signals this affects
All twelve, through the confidence attached to each. The scorecard asks “how do you know?” on every answer for exactly this reason.
Related: Your research expired the day you shipped · What a scorecard cannot tell you