The say-do gap
Your users say it is fine.
Your data says otherwise.
RUCF measures what your product costs users, what they believe it costs them, and the distance between the two.
Silent friction
Behaviour is poor and nobody is complaining. Your users have stopped noticing what this costs them, which means it is invisible to every feedback channel you own.
Composite · illustrative
Score, then confidence. Drag the dot: the reading changes, the arithmetic stays published.
The problem
Every UX framework asks the wrong question
“How good is this product” produces a number nobody has to act on. Rate a checkout 3.4 out of 5 and you have described a feeling with a decimal point attached.
There are five arrangements of the same four dimensions in circulation. Usability, engagement, accessibility, satisfaction. They differ in wording and not much else, and none of them prescribes a different action depending on the answer.
RUCF asks three answerable questions instead. What did the user come to do. What did it cost them. What do they believe happened. Compare those and you get gaps, and the direction of a gap decides the treatment.
Same score. Opposite treatments.
Behaviour and reported experience rarely agree. Most research practice treats that disagreement as noise to be resolved. RUCF treats it as the finding, because the direction of the disagreement decides what you should do.
Hatch · behaviour poor, perception good
Silent friction
Users have absorbed the cost, or they blame themselves for it. No survey will surface this and no support ticket will name it. It is also the cheapest quadrant to fix, because nobody is attached to a design they never noticed.
Treatment. Remove the step outright; explanations and training only preserve it.
Dotted · behaviour good, perception poor
Phantom friction
Tasks complete at reasonable cost and people still call it confusing or slow or untrustworthy. This is the most commonly misdiagnosed state in the discipline, and rebuilding the interaction is the standard wrong answer.
Treatment. Change what the product says, not what it does.
Cross-hatch · both poor
Loud friction
Everyone already knows. Its information value is close to zero and its political value is high, which is why most research budgets end up here.
Treatment. Fix it, and recognise you are paying for confirmation rather than discovery.
Solid · both good
Flow
Behaviour and perception agree, and they agree in the direction you want. A quadrant to defend rather than invest in.
Treatment. Record it as a regression baseline and score every release against it.
Confidence
We never report a score on its own
Measured. You measured this, recently, from behaviour. Act on it.
Guessed. Same score, different life. You have a survey from March and some strong opinions. Go and look.
Observed behaviour enters at 1.00. Reported experience at 0.70. Anything inferred, including AI classification of tickets or session replays, at 0.50. Expert judgment, including a senior designer’s heuristic review, at 0.30. That last figure is the correct weight. Everyone in the field knows it, and no other framework says so numerically.
Confidence also decays, and shipping is what triggers it rather than the calendar. Each release touching a measured surface multiplies its confidence by 0.8, because evidence describes a product that no longer exists. The loop closes itself: shipping a fix lowers confidence, which pushes that surface back up the measurement queue, so re-measurement follows without anyone needing to remember it.
Distributions
We do not report averages
An average experience is a fiction. Most products work well for the majority and badly for a minority, and that minority is where churn, complaints, support cost and legal exposure all live.
4.2
P50: the median session
1.8
P10: the worst tenth
2.4
Tail gap: P50 minus P10
RUCF reports the median and the tenth percentile. If they sit far apart you have not built a product with an accessibility issue. You have built a product for people who resemble the team that made it. Performance engineering stopped reporting mean latency around fifteen years ago for this reason. UX measurement has not caught up.
Six layers, one cycle
- 01
Measures. Intent, friction, perception, captured separately and never allowed to contaminate each other.
- 02
Gaps. Effectiveness, awareness, expectation. Each one a comparison between two measures.
- 03
Diagnosis. Which of the four states each surface is in.
- 04
Confidence. Evidence class, decaying on deploy.
- 05
Priority. Recoverable cost times confidence, divided by effort.
- 06
Loop. Ship the treatment, expire the evidence, measure again.
Limits
What it does not do
RUCF structures user research rather than replacing it. It will not tell you what to build. It will not catch a broken business model or a product nobody wants. It is a poor fit below a few hundred active users, where the sample is too thin for the commitment signals to mean anything. It measures the cost of an experience and the accuracy of your beliefs about it. That is the whole scope.
Get your number
Start with the free scorecard. If the result raises questions, book a session and we will run the full method against real behavioural data.