The say-do gap

Your users say it is fine.
Your data says otherwise.

RUCF measures what your product costs users, what they believe it costs them, and the distance between the two.

Silent friction

Behaviour is poor and nobody is complaining. Your users have stopped noticing what this costs them, which means it is invisible to every feedback channel you own.

Composite · illustrative

2.4@0.62

Score, then confidence. Drag the dot: the reading changes, the arithmetic stays published.

The problem

Every UX framework asks the wrong question

“How good is this product” produces a number nobody has to act on. Rate a checkout 3.4 out of 5 and you have described a feeling with a decimal point attached.

There are five arrangements of the same four dimensions in circulation. Usability, engagement, accessibility, satisfaction. They differ in wording and not much else, and none of them prescribes a different action depending on the answer.

RUCF asks three answerable questions instead. What did the user come to do. What did it cost them. What do they believe happened. Compare those and you get gaps, and the direction of a gap decides the treatment.

Same score. Opposite treatments.

Behaviour and reported experience rarely agree. Most research practice treats that disagreement as noise to be resolved. RUCF treats it as the finding, because the direction of the disagreement decides what you should do.

Hatch · behaviour poor, perception good

Silent friction

Users have absorbed the cost, or they blame themselves for it. No survey will surface this and no support ticket will name it. It is also the cheapest quadrant to fix, because nobody is attached to a design they never noticed.

Treatment. Remove the step outright; explanations and training only preserve it.

Silent friction in full

Dotted · behaviour good, perception poor

Phantom friction

Tasks complete at reasonable cost and people still call it confusing or slow or untrustworthy. This is the most commonly misdiagnosed state in the discipline, and rebuilding the interaction is the standard wrong answer.

Treatment. Change what the product says, not what it does.

Phantom friction in full

Cross-hatch · both poor

Loud friction

Everyone already knows. Its information value is close to zero and its political value is high, which is why most research budgets end up here.

Treatment. Fix it, and recognise you are paying for confirmation rather than discovery.

Loud friction in full

Solid · both good

Flow

Behaviour and perception agree, and they agree in the direction you want. A quadrant to defend rather than invest in.

Treatment. Record it as a regression baseline and score every release against it.

Flow in full

State 01 / 04

The say-do matrix

FLOW PHANTOM FRICTION SILENT FRICTION LOUD FRICTION FRICTION LOW HIGH USERS SAY FINE USERS COMPLAIN

Composite score

2.4@0.71

identical in all four states

Correct treatment

Remove the step outright

One composite score, four positions, four incompatible responses. This is the whole argument for measuring the direction of the disagreement rather than averaging it away.

Confidence

We never report a score on its own

3.8@0.91

Measured. You measured this, recently, from behaviour. Act on it.

3.8@0.34

Guessed. Same score, different life. You have a survey from March and some strong opinions. Go and look.

Observed · 1.00 Reported · 0.70 Inferred · 0.50 Assumed · 0.30

Observed behaviour enters at 1.00. Reported experience at 0.70. Anything inferred, including AI classification of tickets or session replays, at 0.50. Expert judgment, including a senior designer’s heuristic review, at 0.30. That last figure is the correct weight. Everyone in the field knows it, and no other framework says so numerically.

Confidence also decays, and shipping is what triggers it rather than the calendar. Each release touching a measured surface multiplies its confidence by 0.8, because evidence describes a product that no longer exists. The loop closes itself: shipping a fix lowers confidence, which pushes that surface back up the measurement queue, so re-measurement follows without anyone needing to remember it.

Distributions

We do not report averages

An average experience is a fiction. Most products work well for the majority and badly for a minority, and that minority is where churn, complaints, support cost and legal exposure all live.

One release, three readings

4.2

P50: the median session

1.8

P10: the worst tenth

2.4

Tail gap: P50 minus P10

RUCF reports the median and the tenth percentile. If they sit far apart you have not built a product with an accessibility issue. You have built a product for people who resemble the team that made it. Performance engineering stopped reporting mean latency around fifteen years ago for this reason. UX measurement has not caught up.

Six layers, one cycle

One surface, carried through the whole model. Each layer takes the output of the last and adds the thing the one before it could not tell you. Keep scrolling and watch the same reading change shape six times.

Layer 01

Measures

Intent, friction and perception, captured through separate channels and never allowed to inform each other while they are being collected.

Layer 02

Gaps

The three readings become the corners of a triangle, and each edge is a gap. Effectiveness, expectation, and the awareness gap along the base that nothing else reports.

Layer 03

Diagnosis

The awareness gap collapses the three readings into a single position on the matrix, and the position prescribes the treatment. This surface landed in silent friction.

Layer 04

Confidence

The position carries the weight of the evidence behind it, and that weight falls by a fifth on every release that touches the surface. Three ships later it is worth half.

Layer 05

Priority

Recoverable cost times confidence, divided by effort. The smaller finding you observed outranks the larger one you assumed, which is the whole reason confidence is in the formula.

Layer 06

Loop

Shipping the treatment expires the evidence behind it, which raises the surface back up the queue on its own. The measurement schedules itself.

Layer 01 / 06

Three measures

INTENT FRICTION PERCEPTION THREE CHANNELS, NEVER MIXED each carries its own evidence class INTENT FRICTION PERCEPTION AWARENESS GAP EFFECTIVENESS EXPECTATION FLOW PHANTOM SILENT FRICTION LOUD LOW HIGH SAY FINE COMPLAIN 1.00 0.80 0.64 0.51 MEASURED SHIP 1 SHIP 2 SHIP 3 ×0.8 PER RELEASE CHECKOUT £41k @ 0.30 4.1 ACCOUNT SETUP £18k @ 1.00 9.0 OBSERVED BEATS ASSUMED COST × CONFIDENCE ÷ EFFORT MEASURE DIAGNOSE TREAT SHIP EVIDENCE EXPIRES · ×0.8 THE SURFACE RETURNS TO THE QUEUE

The surface

Checkout

the same one throughout

Output of this layer

three readings, three evidence classes

The assembled model

What the six layers produce together

01 MEASURES 3 02 GAPS 3 03 STATE Silent 04 CONFIDENCE 0.86 05 PRIORITY #1 06 LOOP Q3

Checkout costs 2.4, users call it fine, remove the step, re-measure in Q3.

The output is not a rating. It is a sentence with a treatment and an expiry date attached, which is the difference between a framework and a taxonomy.

Read the framework in full

Limits

What it does not do

RUCF structures user research rather than replacing it. It will not tell you what to build. It will not catch a broken business model or a product nobody wants. It is a poor fit below a few hundred active users, where the sample is too thin for the commitment signals to mean anything. It measures the cost of an experience and the accuracy of your beliefs about it. That is the whole scope.

Get your number

Start with the free scorecard. If the result raises questions, book a session and we will run the full method against real behavioural data.