Published in full · CC BY 4.0

The framework

This page describes what RUCF is. If you want to know how an engagement runs week by week, read the methodology instead.

RUCF is published in full and free to use under CC BY 4.0. You do not need us to run it. The model is documented here in enough detail that a competent product team can apply it alone, and we would rather you did that than take anything here on trust.

Version 2.0 · updated August 2026 · changelog

The argument behind the framework

Why divergence beats agreement, with the sources: read the manifesto, the full sourced case for the model on this page.

Gaps, not qualities

RUCF measures gaps, not qualities. A gap is the distance between what a user intended, what the product cost them, and what they believe happened. Every gap carries a cost, a confidence and an expiry date.

Layer 1: three measures

Intent. What the user came to do, in their words. A job with a success condition attached rather than a persona. “Find out whether this covers my situation before I hand over a card number.”

Friction. What it cost them, in four countable currencies: time to outcome, attempts before success, assistance sought, and abandonment at each decision point.

Perception. What they believe happened. Effort felt, confidence in the result, willingness to repeat.

The discipline that makes this work: friction is measured only from behaviour, perception only from report, and the two are never allowed to inform each other during collection. If they contaminate, everything downstream collapses.

Fig. 01

Layer 1 · three measures, three channels

INTENT FRICTION PERCEPTION the job, stated with a success condition TIME TRIES HELP EXIT effort felt confidence in result would repeat ELICITED BEHAVIOUR ONLY REPORT ONLY
Friction is counted in four currencies from behaviour alone. Perception is reported on three scales and never checked against the friction data while it is being collected. The dashed walls are the method: if the channels contaminate each other, every gap downstream is measuring itself.

Layer 2: three gaps

Effectiveness gap. Intent against friction. Did they get what they came for, at what cost. Every existing framework measures this, and it is the least interesting of the three.

Awareness gap. Friction against perception. Does the user know what the product cost them. This is where RUCF earns its existence.

Expectation gap. Intent against perception. Did they get what they thought they would get. This is the gap that produces churn at products with excellent usability. Nothing is broken. It was not the thing they were told it was.

Fig. 02

Layer 2 · three gaps between three measures

INTENT FRICTION PERCEPTION EFFECTIVENESS GAP EXPECTATION GAP AWARENESS GAP measured everywhere churn at good usability
Each edge is a comparison between two measures. The effectiveness gap is the one every existing framework already reports. The awareness gap along the base is the one that resolves into the four states, and it is the reason the model exists.

Layer 3: the say-do matrix

The awareness gap resolves into four states with four different correct treatments. Applying the wrong one is worse than doing nothing, because it spends the budget the right one needed.

Fig. 03

Layer 3 · four states, four treatments

Flow DEFEND AS BASELINE Phantom friction CHANGE WHAT IT SAYS Silent friction REMOVE THE STEP Loud friction FIX IT FRICTION LOW HIGH USERS SAY FINE USERS COMPLAIN PERCEPTION
The level of friction does not decide the treatment. The direction of the disagreement does. Two surfaces can carry an identical composite score and sit in opposite quadrants, and the treatment that repairs one wastes the budget on the other.

Read the full matrix

Layer 4: confidence

Evidence classWhat it isWeight
ObservedTelemetry, session data, moderated task tests1.00
ReportedSurveys, interviews, support tickets, reviews0.70
InferredAnalytics proxies, AI classification of text or replays0.50
AssumedTeam judgment, heuristic review, expert opinion0.30

Two of these are deliberately uncomfortable.

Expert opinion is the weakest evidence in the system. Anything a model classified enters at 0.50 and stays there until a human validates a sample, at which point it is promoted to the class of the validation method. A framework that discounts AI evidence in its own arithmetic is more trustworthy than one that markets AI as a feature.

Confidence then decays on deploy rather than on the calendar. confidence × 0.8 per release touching the measured surface, plus a slow time half-life of twelve months for reported evidence and eighteen for observed, because user expectations drift even when your product does not.

Fig. 04

Layer 4 · what evidence is worth, and for how long

CLASS WEIGHT DECAY ON DEPLOY OBSERVED 1.00 REPORTED 0.70 INFERRED 0.50 ASSUMED 0.30 expert judgment, drawn hollow: the weakest class in the system 1.00 0.80 0.64 0.51 MEASURED SHIP 1 SHIP 2 SHIP 3 ×0.8 PER RELEASE
A score is never reported without the second number. Three releases after the measurement, an observed reading is worth about half of what it was, and that decay is what pushes the surface back up the measurement queue without anyone having to remember it.

Layer 5: priority

priority = recoverable cost × confidence ÷ effort

Recoverable cost is stated in your currency: abandonment times traffic times value per completion, or support contacts times cost per contact, or seconds lost times sessions for an internal tool.

Confidence enters the priority calculation directly, so a large cost you are unsure about ranks below a smaller cost you have observed. Every prioritisation model that omits confidence gets used to justify expensive work on a guess.

Fig. 05

Layer 5 · the cheaper finding outranks the bigger guess

SURFACE COST CONF. EFFORT PRIORITY Checkout £41k 0.30 3 4.1 Account setup £18k 1.00 2 9.0 RANKS FIRST ON LESS MONEY observed beats assumed priority = recoverable cost × confidence ÷ effort
Checkout looks like the obvious target until the confidence figure enters the arithmetic. A prioritisation model without a confidence term will always rank the largest guess above the smallest certainty, which is how expensive work gets justified by an opinion.

Layer 6: the loop

Measure, diagnose, treat, ship, expire, measure again. Shipping expires the evidence, which raises that surface back up the queue automatically.

Fig. 06

Layer 6 · the loop closes itself

01 02 03 04 05 MEASURE DIAGNOSE TREAT SHIP EXPIRE SHIPPING EXPIRES THE EVIDENCE · ×0.8
The return path is the mechanism, not a diagram convention. Treating a surface and shipping the fix lowers the confidence attached to it, which raises it back up the priority queue on its own. Nobody has to remember to re-measure.

How the loop runs

Accessibility, restructured

Conformance is a gate. WCAG 2.2 AA is a pass or fail against a published standard with legal force in the EU. Averaging it into a composite lets an attractive product post a respectable number while failing a legal requirement, so RUCF removes it from the score entirely. Fail conformance and RUCF returns a remediation list instead of an index.

Accessibility experience lives in the tail of every dimension rather than in a dimension of its own. A product can conform perfectly and still be miserable to use with a screen reader, one hand, or a poor connection, and that misery lives in the tenth percentile of usability, of commitment, of everything.

What RUCF does not do

It structures user research rather than replacing it. It does not tell you what to build. It will not catch a broken business model. It is a poor fit below a few hundred active users. It measures the cost of an experience and the accuracy of your beliefs about that cost.

Prior art

None of the mechanisms here are unprecedented. Assembling them is the contribution.

IdeaSourceWhat RUCF does with it
Stated against revealed preferenceEconomics, Samuelson 1938Becomes the primary diagnostic axis instead of a caveat
The attitude-behaviour gapSocial psychology, LaPiere 1934 onwardFormalised into four states with distinct treatments
Evidence quality weightingGRADE, evidence-based medicineAdapted to UX evidence classes with numeric weights
Tail percentile reportingPerformance engineering, p99 latencyApplied to experience as P50 and P10
Effort-weighted prioritisationRICE and ICEExtended with a confidence multiplier
Task-based usability measurementISO 9241-11, NielsenKept inside the friction measure
Conformance as a standardWCAG 2.2, EU Accessibility ActRemoved from scoring, promoted to a gate

Licence

CC BY 4.0. Use it, teach it, adapt it, run it for clients. Attribute it to RUCF and link back. You do not need permission and you do not owe anything.

Changelog

v2.0, August 2026. Complete rebuild. v1 scored four attributes and averaged them, which made it a taxonomy rather than a framework. v2 measures gaps, introduces the say-do matrix, confidence weighting, deploy decay and tail reporting. The twelve signals survive as the friction inventory. Accessibility moved from a scored dimension to a gate plus a tail.

v1.0, 2025. Four dimensions, twelve signals, composite average. Retired.

See which state your product is in