Published in full · CC BY 4.0
The framework
This page describes what RUCF is. If you want to know how an engagement runs week by week, read the methodology instead.
RUCF is published in full and free to use under CC BY 4.0. You do not need us to run it. The model is documented here in enough detail that a competent product team can apply it alone, and we would rather you did that than take anything here on trust.
Version 2.0 · updated August 2026 · changelog
The argument behind the framework
Why divergence beats agreement, with the sources: read the manifesto — the full sourced case for the model on this page.
The core idea
Gaps, not qualities
RUCF measures gaps, not qualities. A gap is the distance between what a user intended, what the product cost them, and what they believe happened. Every gap carries a cost, a confidence and an expiry date.
Layer 1: three measures
Intent. What the user came to do, in their words. A job with a success condition attached, not a persona. “Find out whether this covers my situation before I hand over a card number.”
Friction. What it cost them, in four countable currencies: time to outcome, attempts before success, assistance sought, and abandonment at each decision point.
Perception. What they believe happened. Effort felt, confidence in the result, willingness to repeat.
The discipline that makes this work: friction is measured only from behaviour, perception only from report, and the two are never allowed to inform each other during collection. If they contaminate, everything downstream collapses.
Layer 2: three gaps
Effectiveness gap. Intent against friction. Did they get what they came for, at what cost. Every existing framework measures this, and it is the least interesting of the three.
Awareness gap. Friction against perception. Does the user know what the product cost them. This is where RUCF earns its existence.
Expectation gap. Intent against perception. Did they get what they thought they would get. This is the gap that produces churn at products with excellent usability. Nothing is broken. It simply was not the thing they were told it was.
Layer 3: the say–do matrix
The awareness gap resolves into four states with four different correct treatments. Applying the wrong one is worse than doing nothing, because it spends the budget the right one needed.
Layer 4: confidence
| Evidence class | What it is | Weight |
|---|---|---|
| Observed | Telemetry, session data, moderated task tests | 1.00 |
| Reported | Surveys, interviews, support tickets, reviews | 0.70 |
| Inferred | Analytics proxies, AI classification of text or replays | 0.50 |
| Assumed | Team judgment, heuristic review, expert opinion | 0.30 |
Two of these are deliberately uncomfortable.
Expert opinion is the weakest evidence in the system. Anything a model classified enters at 0.50 and stays there until a human validates a sample, at which point it is promoted to the class of the validation method. A framework that discounts AI evidence in its own arithmetic is more trustworthy than one that markets AI as a feature.
Confidence then decays on deploy, not on the calendar. confidence × 0.8 per release touching the measured surface, plus a slow time half-life of twelve months for reported evidence and eighteen for observed, because user expectations drift even when your product does not.
Layer 5: priority
priority = recoverable cost × confidence ÷ effort
Recoverable cost is stated in your currency: abandonment times traffic times value per completion, or support contacts times cost per contact, or seconds lost times sessions for an internal tool.
Confidence enters the priority calculation directly, so a large cost you are unsure about ranks below a smaller cost you have observed. Every prioritisation model that omits confidence gets used to justify expensive work on a guess.
Layer 6: the loop
Measure, diagnose, treat, ship, expire, measure again. Shipping expires the evidence, which raises that surface back up the queue automatically.
Accessibility, restructured
Conformance is a gate, not a score. WCAG 2.2 AA is a pass or fail against a published standard with legal force in the EU. Averaging it into a composite lets an attractive product post a respectable number while failing a legal requirement, so RUCF removes it from the score entirely. Fail conformance and you do not get an index, you get a remediation list.
Accessibility experience is not a dimension either. It is the tail of every dimension. A product can conform perfectly and still be miserable to use with a screen reader, one hand, or a poor connection, and that misery lives in the tenth percentile of usability, of commitment, of everything.
What RUCF does not do
It does not replace user research, it structures it. It does not tell you what to build. It will not catch a broken business model. It is a poor fit below a few hundred active users. It measures the cost of an experience and the accuracy of your beliefs about that cost.
Prior art
None of the mechanisms here are unprecedented. Assembling them is the contribution.
| Idea | Source | What RUCF does with it |
|---|---|---|
| Stated against revealed preference | Economics, Samuelson 1938 | Becomes the primary diagnostic axis instead of a caveat |
| The attitude–behaviour gap | Social psychology, LaPiere 1934 onward | Formalised into four states with distinct treatments |
| Evidence quality weighting | GRADE, evidence-based medicine | Adapted to UX evidence classes with numeric weights |
| Tail percentile reporting | Performance engineering, p99 latency | Applied to experience as P50 and P10 |
| Effort-weighted prioritisation | RICE and ICE | Extended with a confidence multiplier |
| Task-based usability measurement | ISO 9241-11, Nielsen | Kept inside the friction measure |
| Conformance as a standard | WCAG 2.2, EU Accessibility Act | Removed from scoring, promoted to a gate |
Licence
CC BY 4.0. Use it, teach it, adapt it, run it for clients. Attribute it to RUCF and link back. You do not need permission and you do not owe anything.
Changelog
v2.0, August 2026. Complete rebuild. v1 scored four attributes and averaged them, which made it a taxonomy rather than a framework. v2 measures gaps, introduces the say–do matrix, confidence weighting, deploy decay and tail reporting. The twelve signals survive as the friction inventory. Accessibility moved from a scored dimension to a gate plus a tail.
v1.0, 2025. Four dimensions, twelve signals, composite average. Retired.