Manifesto · version 2.0 · August 2026 · CC BY 4.0
Measure the divergence, not the agreement
A manifesto for the Real User Cost Framework. Iordanis Passas.
Save as PDF · Cite: Passas, I. (2026). Measure the divergence, not the agreement (v2.0). rucframework.com/manifesto/
Abstract
Contemporary user experience measurement collapses attitudinal and behavioural evidence into composite scores. This paper argues that the collapse destroys the most diagnostically valuable signal available: the direction and magnitude of divergence between what users do and what users report. Drawing on eight decades of attitude–behaviour research, on evidence-grading practice from clinical medicine, and on distribution reporting from performance engineering, it proposes a framework in which divergence is the primary diagnostic axis, every measurement carries an evidence-quality weight, that weight decays on software release rather than on the calendar, and results are reported as distributions rather than means. The framework is falsifiable, and section 9 states the conditions under which it should be rejected.
1 · The problem: measurement that cannot be wrong
A usability score of 3.4 out of 5 cannot be wrong. It cannot be wrong because it is not a claim about the world. It is a summary of judgments, and the judgments were formed by the same people who assign the number.
This is not a minor complaint about calibration. It is a category problem. A measurement instrument earns its status by being capable of returning a result its operator did not expect and does not want. An instrument that returns whatever the operator already believed is not measuring, it is recording.
The dominant models in the field share this property to varying degrees. ISO 9241-11 defines usability as effectiveness, efficiency and satisfaction. Nielsen and Molich's heuristic evaluation asks experts to identify violations against ten principles. The System Usability Scale asks users ten Likert items and produces a number benchmarked around 68. Google's HEART framework organises happiness, engagement, adoption, retention and task success into a goals-signals-metrics structure.
Each is useful. None of them prescribes a different action depending on the answer, beyond “the low one needs attention”. And several of them accept attitudinal self-report as a proxy for experienced cost, which section 2 argues is untenable.
The claim of this paper is narrow and specific. It is not that these instruments are worthless. It is that they are constructed to produce agreement, and agreement is the least informative thing a measurement can produce.
2 · Self-report and behaviour diverge systematically
2.1 The finding is old
LaPiere travelled the United States with a young Chinese couple in the early 1930s, visiting several hundred establishments. Nearly all served them. Months later he wrote to the same establishments asking whether they would accept Chinese guests, and the overwhelming majority said they would not. Stated attitude and observed behaviour pointed in opposite directions with near-total consistency.
Wicker reviewed the accumulated literature three decades later and concluded that attitudes were only weakly related to overt behaviour, with correlations rarely exceeding 0.30. Ajzen and Fishbein subsequently established the correspondence principle: attitude measures predict behaviour only to the degree that they match the behaviour in action, target, context and time. That refinement rescued attitude research. It did not rescue the practice of asking a general satisfaction question and treating the answer as a proxy for what a specific interaction cost.
2.2 The mechanism
Nisbett and Wilson's review remains the clearest statement of why this happens. People have little direct introspective access to their higher-order cognitive processes. When asked why they did something, they do not retrieve the cause. They generate a plausible account from available theories about what ought to have caused it, and they report that account with confidence indistinguishable from genuine recall.
This has an immediate consequence for research practice. A user who spent eleven seconds and three attempts on a step is not withholding that from you. They do not have it. They have an impression, assembled after the fact, from whatever was salient.
2.3 What is salient is not what was costly
Kahneman and colleagues demonstrated that retrospective evaluations of an experience are dominated by its peak and its ending, and are largely insensitive to duration. Redelmeier and Kahneman found that patients rated a longer procedure with a less painful ending as better than a shorter procedure with a sharper finish, despite the longer one containing strictly more total discomfort.
Translated into product terms: a checkout that is painful throughout but ends with a satisfying confirmation will be remembered better than a checkout that is smoother but ends flatly. The remembered experience and the incurred cost are different quantities, produced by different processes, and there is no reason to expect them to agree.
2.4 Contemporary quantification
The most direct modern measurement of the gap comes from Parry and colleagues, who conducted a pre-registered meta-analysis of studies that captured both self-reported and log-based digital media use. Across 106 effect sizes, self-reported use correlated with logged use at r = 0.38 (95% CI 0.33 to 0.42). Of 49 comparisons of mean estimates, only three, about 6%, produced an accurate reflection of logged use.
An r of 0.38 means self-report accounts for roughly 14% of the variance in actual behaviour. The remaining 86% is not noise around a true signal. It is a second, different quantity.
2.5 The inference
The literature does not support treating the divergence as measurement error to be minimised. It supports treating it as a stable, structured property of human self-knowledge.
If that is right, then a framework that averages the two quantities together is not reducing error. It is deliberately discarding the second variable.
3 · Why averaging destroys the diagnostic signal
State the problem formally. Let a product surface have a behavioural cost measure B and a reported experience measure P, both normalised to [0, 100].
Conventional practice reports some composite C = f(B, P), typically a weighted mean. The mapping is many-to-one, so C is not invertible: from C alone the pair (B, P) cannot be recovered.
Consider two surfaces:
Identical composites. Now consider what each state implies.
In α, users are paying a substantial cost and reporting satisfaction. Whatever is wrong is invisible to every feedback channel the organisation operates, because feedback channels sample P. No survey will surface it, no support ticket will name it, no interview will produce it unprompted.
In β, users are paying little and reporting dissatisfaction. The interaction functions. The belief about the interaction does not. Reconstructing the interaction, the standard response to a poor experience score, addresses a variable that is already healthy.
The correct interventions are not merely different in degree. They are different in kind, and applying α's treatment to β is expensive and inert.
This is the argument in full. Any framework reporting only
Ccannot distinguish α from β, and therefore cannot prescribe. The information required to act was present in the inputs and was destroyed by the aggregation function. RetainingD = P − BalongsideCcosts nothing and recovers the entire diagnosis.
4 · The framework
4.1 Three measures
RUCF captures three quantities per user journey and forbids their contamination during collection.
Intent I — the job the user came to accomplish, with a success condition, stated in the user's terms.
Friction B — the cost incurred, expressed in four countable currencies: time to outcome, attempts before success, assistance sought, and abandonment at each decision point. Sourced exclusively from observation.
Perception P — the believed experience: effort felt, confidence in the outcome, willingness to repeat. Sourced exclusively from report.
The separation is methodological, not stylistic. Instruments that ask users to estimate their own behaviour, or that show researchers behavioural data before they gather perception data, contaminate the two variables and make the divergence uncomputable.
4.2 Three gaps
The effectiveness gap is what existing frameworks measure. The awareness gap is the contribution of this framework. The expectation gap explains churn at products with excellent usability, where nothing is broken and the product was simply not the thing the user was told it was.
4.3 The say–do matrix
The awareness gap resolves into four states plus a dead band. With thresholds HI = 58, LO = 42:
| State | Condition | Interpretation | Treatment |
|---|---|---|---|
| Flow | B ≥ HI, P ≥ HI | Working and believed to work | Record as regression baseline |
| Silent friction | B < LO, P ≥ HI | Cost incurred, not perceived | Remove the step. Do not explain it. |
| Phantom friction | B ≥ HI, P < LO | Cost not incurred, perceived | Change what the product says, not what it does |
| Loud friction | B < LO, P < LO | Cost incurred and perceived | Fix it, and note you are buying confirmation |
| Watch | otherwise | Inside the dead band | Instrument before acting |
The dead band exists because a threshold without one produces diagnoses that flip on a single response, which is a reliability failure rather than a finding.
4.4 Silent friction as the principal object of interest
Three mechanisms produce the silent state, and each has a literature.
Normalisation. Repeated exposure moves the cost to baseline. There is no comparison class, so “acceptable” is defined by whatever is present.
Self-attribution. Norman's central argument in The Design of Everyday Things is that users blame themselves for failures caused by design, and that the label “human error” typically conceals a design defect. A user who believes the difficulty is theirs will not report it as a product problem.
Introspective unavailability. Per section 2.2, the user cannot report a cost they did not consciously register.
The methodological consequence is the sharpest claim in this paper: silent friction is undetectable by any perception-sampling method. Surveys, interviews, feedback widgets and satisfaction instruments all sample
P. WherePis precisely the variable that is wrong, the instrument returns a clean result. This is not a failure of practitioner skill. It is a structural property of asking questions, and the only remedy is to measureBindependently and compare.
5 · Evidence quality as a first-class variable
5.1 The borrowing
Clinical medicine does not treat all evidence as equivalent. The GRADE framework rates the certainty of evidence across studies and makes that rating explicit in every recommendation, so that a recommendation and the confidence in it travel together.
UX measurement has no equivalent. A score derived from instrumented telemetry and a score derived from a workshop are reported in the same units, in the same font, with the same authority.
5.2 The weights
| Class | Source | Weight |
|---|---|---|
| Observed | Telemetry, session data, moderated task testing | 1.00 |
| Reported | Surveys, interviews, support records, reviews | 0.70 |
| Inferred | Analytics proxies, model classification of text or replay | 0.50 |
| Assumed | Expert judgment, heuristic review, team consensus | 0.30 |
5.3 Justifying 0.30 for expert judgment
This is the least comfortable number in the framework and it is the best supported.
The evaluator effect is the finding that different evaluators examining the same system with the same method report substantially different problem sets. Hertzum and Jacobsen reviewed eleven studies and found average agreement between evaluators ranging from 5% to 65%.
The primary studies are more pointed still. Jacobsen, Hertzum and John had four evaluators independently analyse the same four videotaped sessions. Of 93 problems detected in total, only 20% were detected by all four, and 46% were detected by a single evaluator alone. When each selected their ten most severe problems, no problem appeared on all four lists.
Hertzum, Jacobsen and Molich had eleven usability specialists individually inspect one website. Average pairwise overlap in reported problems was 9%. The specialists nonetheless believed they were largely in agreement, interpreting their disparate observations as corroboration.
That last result deserves emphasis, because it is the evaluator-effect analogue of the say–do gap. Eleven experts overlapped by 9% and experienced that as consensus. Their perception of their own reliability diverged from their measured reliability in exactly the direction this framework predicts for users.
A later replication found participants reporting on average 33% of the problems identified across the full group, with 24% to 30% of multiply-reported problems rated critical by one evaluator and minor by another.
A weight of 0.30 for expert judgment is therefore not a provocation. Given inter-rater agreement measured between 9% and 33% in controlled conditions, it is arguably generous.
This applies to the author. Nothing in the framework's provenance exempts its author's judgment from its own weighting.
5.4 Model classification at 0.50
Machine classification of unstructured evidence, including support records, session replays and interview transcripts, enters at 0.50 and remains there until a human validates a representative sample, at which point it is promoted to the class of the validation method.
The reasoning is that current classification systems are strong at volume and unreliable at nuance, and that their failure mode is confident rather than conspicuous. A framework whose arithmetic discounts machine evidence is more defensible than one that markets it, and the promotion path means the discount is an incentive to validate rather than a prohibition on use.
6 · Temporal validity: evidence expires on release
6.1 The problem with calendar decay
Research findings are conventionally treated as valid until someone judges them stale. In practice the judgment is rarely made, and repositories accumulate findings about interfaces that no longer exist.
Calendar-based expiry is the obvious correction and it is the wrong one. A finding about a surface untouched for two years may remain perfectly valid. A finding about a surface redesigned last Tuesday is void regardless of how recently it was gathered.
6.2 Change, not time
Lehman's laws of software evolution establish that a system in use undergoes continuous change and that its structure degrades unless work is done to prevent it. The relevant implication here is simpler than the laws themselves: the artefact your evidence describes is not the artefact currently in production.
RUCF therefore decays confidence on deployment:
Three releases reduce confidence to roughly half regardless of recency. The slow time term remains because user expectations drift even where the product does not.
6.3 The self-closing loop
This produces a property worth stating separately. Shipping a treatment lowers the confidence attached to the treated surface, which raises that surface in the measurement queue, which prompts re-measurement.
The loop closes through arithmetic rather than discipline. Frameworks that depend on organisational virtue to re-measure do not get re-measured.
Implementation requires only that signals are tagged to surfaces and surfaces are named in release notes. It is automatable from a deployment pipeline.
7 · Distributions, not central tendency
7.1 The borrowed correction
Large-scale systems engineering abandoned mean latency as a primary metric. Dean and Barroso's account of tail latency shows that in systems where a user-facing operation depends on many components, rare slow responses dominate the experience, and that mean response time conceals precisely the behaviour that determines whether the system is usable.
The structural analogy to experience measurement is exact. A product with a mean score of 4.2 may be excellent for most users and unusable for a minority, and the mean is constructed to hide that.
7.2 What RUCF reports
Every measure is reported at the median and the tenth percentile, with the distance between them named as the tail gap.
A product at P50 4.2 and P10 1.8 has not been built for a general population. It has been built for a population resembling the team that built it, and the tail gap quantifies the resemblance.
7.3 Accessibility, reclassified
Two changes follow, and both are structural rather than presentational.
Conformance becomes a gate. WCAG 2.2 defines testable success criteria at levels A, AA and AAA, and Directive (EU) 2019/882 gives conformance legal force for in-scope products in the European Union. A binary legal requirement averaged into a composite permits an otherwise strong product to report a respectable figure while failing a statutory obligation. RUCF removes conformance from the composite entirely: a product failing conformance receives a remediation list, not an index.
Accessibility experience becomes the tail. A product can conform fully and remain difficult to use with a screen reader, with one hand, on a constrained device or under time pressure. That difficulty does not occupy a separate dimension. It occupies the tenth percentile of every dimension. Treating accessibility as one dimension among four permits it to be averaged away by the others, which is both an analytical error and, given the directive, an operational risk.
8 · Limitations and threats to validity
A framework that does not publish its own weaknesses is asking to be trusted rather than tested.
8.1 The thresholds are calibrated judgments. HI = 58, LO = 42, the dead band, and the evidence weights are reasoned choices, not derived constants. They are published, versioned and open to challenge. GRADE's certainty ratings share this property and remain useful, but the honest statement is that these numbers await empirical calibration against outcome data.
8.2 The public scorecard is itself self-report. The free tool asks practitioners to report their own behavioural figures, which introduces exactly the reporting bias the framework is built to detect. The mitigation is that it demands an evidence class per answer and returns a low confidence figure with an instrumentation list when the inputs are estimated. The mitigation is partial. A scorecard result is a direction, not a finding, and the tool says so.
8.3 Construct validity is unestablished. Whether the twelve signals collectively measure “the cost of an experience” is a construct-validity claim in the sense of Cronbach and Meehl, and it has not been tested through a nomological network. It is a designed taxonomy, defensible and unvalidated.
8.4 No predictive validation yet exists. The framework asserts that closing gaps in high-priority surfaces produces measurable commercial outcomes. That assertion has not been tested longitudinally. Section 10 proposes the study.
8.5 The framework is vulnerable to its own metric becoming a target. Campbell's law holds that quantitative indicators used for social decision-making are subject to corruption pressure and tend to distort the processes they monitor, a point made independently in Goodhart's observation about regulatory targets. A composite RUCF score adopted as an organisational objective will be gamed, most easily by classifying assumed evidence as observed. The confidence model is a partial defence, because inflating confidence requires claiming measurements that can be audited. It is not a complete one.
8.6 Small populations. Below roughly a few hundred active users the commitment signals lack the sample to be meaningful, and the tail percentile is uncomputable. The framework should not be applied there.
8.7 The say–do gap is not a novel discovery. It is documented across economics and psychology as cited above. The contribution claimed here is operational rather than empirical: assembling a known divergence into a scoring architecture with quadrant-specific treatments, evidence weighting and deployment-triggered expiry. Readers should evaluate that claim and not a stronger one.
9 · Falsifiability
The framework makes claims that can fail. Stating them is the point of publishing rather than selling.
F1. If, across a reasonable sample of products, behavioural and perceptual measures of the same surface correlate above approximately r = 0.8, then divergence is rare, the matrix collapses toward its diagonal, and the framework's central premise is false for product experience even if it holds elsewhere.
F2. If silent-friction surfaces treated by subtraction show no greater behavioural improvement than the same surfaces treated by explanation, the quadrant-specific treatment logic adds nothing and should be discarded.
F3. If evidence gathered before N releases predicts current behaviour as well as evidence gathered after them, deployment decay is unjustified and should be replaced by calendar decay or removed.
F4. If practitioners applying the published criteria independently to the same product diverge more than approximately 0.8 on the composite, the framework has the reliability problem it accuses heuristic evaluation of having, and it should be treated as an aid to thinking rather than an instrument.
F5. If tail gap does not correlate with churn, support cost or complaint volume in any observed dataset, distribution reporting is decoration and the median suffices.
The author's position is that F4 is the most likely of the five to fail, and it is the first that should be tested.
10 · Research agenda
- Threshold calibration. Establish
HIandLOempirically against outcome data rather than by reasoning. - Inter-rater reliability. Run the evaluator-effect protocol on RUCF itself: multiple practitioners, same product, published criteria, measured agreement. Publish the result whatever it shows.
- Treatment efficacy by quadrant. Controlled comparison of subtraction against explanation in silent-friction surfaces.
- Decay validation. Test whether the 0.8-per-release coefficient reflects observed loss of predictive validity.
- Tail gap and commercial outcome. Test the correlation asserted in F5.
- Construct validity. Test whether the twelve signals behave as a coherent construct.
Collaboration on any of these is welcome, including from researchers intending to refute the framework, who are the more useful collaborators.
11 · Position
Three commitments, stated plainly.
Divergence is data. The disagreement between what users do and what users say is not error to be minimised. It is the most diagnostically valuable quantity available, and averaging it away is the field's most consequential habit.
Confidence travels with the number. A measurement reported without its evidential basis is an assertion. This applies to expert judgment, to machine classification, and to this framework's author.
Distributions over means. The median describes a user who may not exist. The tenth percentile describes users who do.
References
- Ajzen, I., & Fishbein, M. (1977). Attitude-behavior relations: A theoretical analysis and review of empirical research. Psychological Bulletin, 84(5), 888–918.
- Brooke, J. (1996). SUS: A quick and dirty usability scale. In P. W. Jordan et al. (Eds.), Usability Evaluation in Industry. Taylor & Francis.
- Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90.
- Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
- Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80.
- European Union. (2019). Directive (EU) 2019/882 on the accessibility requirements for products and services.
- Guyatt, G. H., et al. (2008). GRADE: An emerging consensus on rating quality of evidence and strength of recommendations. BMJ, 336(7650), 924–926.
- Hertzum, M., & Jacobsen, N. E. (2003). The evaluator effect: A chilling fact about usability evaluation methods. International Journal of Human-Computer Interaction, 15(1), 183–204.
- Hertzum, M., Jacobsen, N. E., & Molich, R. (2002). Usability inspections by groups of specialists: Perceived agreement in spite of disparate observations. CHI 2002 Extended Abstracts, 662–663. ACM Press.
- Hertzum, M., Molich, R., & Jacobsen, N. E. (2014). What you get is what you see: Revisiting the evaluator effect in usability tests. Behaviour & Information Technology, 33(2), 143–161.
- ISO. (2018). ISO 9241-11:2018 Ergonomics of human-system interaction, Part 11: Usability: Definitions and concepts.
- Jacobsen, N. E., Hertzum, M., & John, B. E. (1998). The evaluator effect in usability studies: Problem detection and severity judgments. Proceedings of the Human Factors and Ergonomics Society 42nd Annual Meeting, 1336–1340.
- Kahneman, D., Fredrickson, B. L., Schreiber, C. A., & Redelmeier, D. A. (1993). When more pain is preferred to less: Adding a better end. Psychological Science, 4(6), 401–405.
- LaPiere, R. T. (1934). Attitudes vs. actions. Social Forces, 13(2), 230–237.
- Lehman, M. M. (1980). Programs, life cycles, and laws of software evolution. Proceedings of the IEEE, 68(9), 1060–1076.
- Nielsen, J., & Molich, R. (1990). Heuristic evaluation of user interfaces. Proceedings of CHI '90, 249–256. ACM Press.
- Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259.
- Norman, D. A. (2013). The Design of Everyday Things (Revised and expanded ed.). Basic Books.
- Parry, D. A., Davidson, B. I., Sewall, C. J. R., Fisher, J. T., Mieczkowski, H., & Quintana, D. S. (2021). A systematic review and meta-analysis of discrepancies between logged and self-reported digital media use. Nature Human Behaviour, 5(11), 1535–1547.
- Redelmeier, D. A., & Kahneman, D. (1996). Patients' memories of painful medical treatments. Pain, 66(1), 3–8.
- Rodden, K., Hutchinson, H., & Fu, X. (2010). Measuring the user experience on a large scale: User-centered metrics for web applications. Proceedings of CHI 2010. ACM Press.
- Samuelson, P. A. (1938). A note on the pure theory of consumer's behaviour. Economica, 5(17), 61–71.
- Sauro, J., & Lewis, J. R. (2016). Quantifying the User Experience (2nd ed.). Morgan Kaufmann.
- Strathern, M. (1997). “Improving ratings”: Audit in the British University system. European Review, 5(3), 305–321.
- W3C. (2023). Web Content Accessibility Guidelines (WCAG) 2.2. W3C Recommendation.
- Wicker, A. W. (1969). Attitudes versus actions: The relationship of verbal and overt behavioral responses to attitude objects. Journal of Social Issues, 25(4), 41–78.
Cite this document as: Passas, I. (2026). Measure the divergence, not the agreement: A manifesto for the Real User Cost Framework (Version 2.0). https://rucframework.com/manifesto/
If you want to see the framework applied rather than argued: the scorecard.