The say-do gap

Your users say it is fine.
Your data says otherwise.

RUCF measures what your product costs users, what they believe it costs them, and the distance between the two.

Silent friction

Behaviour is poor and nobody is complaining. Your users have stopped noticing what this costs them, which means it is invisible to every feedback channel you own.

Composite · illustrative

2.4@0.62

Score, then confidence. Drag the dot: the reading changes, the arithmetic stays published.

The problem

Every UX framework asks the wrong question

“How good is this product” produces a number nobody has to act on. Rate a checkout 3.4 out of 5 and you have described a feeling with a decimal point attached.

There are five arrangements of the same four dimensions in circulation. Usability, engagement, accessibility, satisfaction. They differ in wording and not much else, and none of them prescribes a different action depending on the answer.

RUCF asks three answerable questions instead. What did the user come to do. What did it cost them. What do they believe happened. Compare those and you get gaps, and the direction of a gap decides the treatment.

Confidence

We never report a score on its own

3.8@0.91

Measured. You measured this, recently, from behaviour. Act on it.

3.8@0.34

Guessed. Same score, different life. You have a survey from March and some strong opinions. Go and look.

Observed · 1.00 Reported · 0.70 Inferred · 0.50 Assumed · 0.30

Observed behaviour enters at 1.00. Reported experience at 0.70. Anything inferred, including AI classification of tickets or session replays, at 0.50. Expert judgment, including a senior designer’s heuristic review, at 0.30. That last figure is the correct weight. Everyone in the field knows it, and no other framework says so numerically.

Confidence also decays, and shipping is what triggers it rather than the calendar. Each release touching a measured surface multiplies its confidence by 0.8, because evidence describes a product that no longer exists. The loop closes itself: shipping a fix lowers confidence, which pushes that surface back up the measurement queue, so re-measurement follows without anyone needing to remember it.

Distributions

We do not report averages

An average experience is a fiction. Most products work well for the majority and badly for a minority, and that minority is where churn, complaints, support cost and legal exposure all live.

4.2

P50: the median session

1.8

P10: the worst tenth

2.4

Tail gap: P50 minus P10

RUCF reports the median and the tenth percentile. If they sit far apart you have not built a product with an accessibility issue. You have built a product for people who resemble the team that made it. Performance engineering stopped reporting mean latency around fifteen years ago for this reason. UX measurement has not caught up.

Six layers, one cycle

  1. 01

    Measures. Intent, friction, perception, captured separately and never allowed to contaminate each other.

  2. 02

    Gaps. Effectiveness, awareness, expectation. Each one a comparison between two measures.

  3. 03

    Diagnosis. Which of the four states each surface is in.

  4. 04

    Confidence. Evidence class, decaying on deploy.

  5. 05

    Priority. Recoverable cost times confidence, divided by effort.

  6. 06

    Loop. Ship the treatment, expire the evidence, measure again.

Read the framework

Limits

What it does not do

RUCF structures user research rather than replacing it. It will not tell you what to build. It will not catch a broken business model or a product nobody wants. It is a poor fit below a few hundred active users, where the sample is too thin for the commitment signals to mean anything. It measures the cost of an experience and the accuracy of your beliefs about it. That is the whole scope.

Get your number

Start with the free scorecard. If the result raises questions, book a session and we will run the full method against real behavioural data.