Methodology

How an engagement runs

The framework is free and documented in full. This page is about what changes when we run it with you: real behavioural data instead of self-assessment, and a baseline you can both defend.

  1. 01

    Baseline

    We score your product independently against the twelve signals, using whatever evidence already exists. You get a scored baseline with the evidence class attached to every signal, and an honest confidence figure that is usually lower than anyone expects.

    Produces: scored baseline, evidence inventory, list of what is currently unmeasurable. Takes: one to two weeks.

    Phase 01

    Twelve signals, graded by what you can prove

    TASK ACCESS BELIEF COMMITMENT KEY observed 1.00 reported 0.70 assumed 0.30 COMPOSITE CONFIDENCE 0.41 most of what you believe about your product is currently opinion
    The first deliverable is usually uncomfortable. Scoring what already exists tends to reveal that three or four of the twelve signals rest on observed behaviour and the rest rest on the team's judgment, which the model weights at 0.30 and says so.
  2. 02

    Research

    Interviews, session analysis and behavioural data. This is where self-assessment gets replaced by observed behaviour, and where the score usually moves, often downward, because guesses are optimistic.

    Produces: revised scores at observed confidence, intent map, friction inventory in your currency. Takes: three to four weeks.

    Phase 02

    Guesses meet behaviour, and usually lose

    ASSUMED SCORE OBSERVED SCORE TASK 4.1 2.8 −1.3 ACCESS 3.6 3.4 −0.2 BELIEF 4.4 2.1 −2.3 COMMITMENT 3.2 3.5 +0.3 CONFIDENCE 0.41 0.86 the number that matters moved most
    Scores generally fall in this phase, because guesses are optimistic and belief is the surface teams misread worst. One surface going up is normal. The real deliverable is the confidence figure, which roughly doubles once the evidence is observed rather than assumed.
  3. 03

    Diagnosis

    Each surface placed in the matrix. Weakest signals ranked by effort against expected movement. Not everything is worth fixing, and this phase is largely about saying so.

    Produces: quadrant map, ranked treatment list with cost and effort per item. Takes: one week.

    Phase 03

    Every surface placed, then ranked

    QUADRANT MAP SAY FINE COMPLAIN CHECKOUT SEARCH SIGN-IN BILLING RANKED TREATMENTS 01 Checkout 02 Search copy 03 Billing 04 Sign-in 05 Empty states NOT WORTH FIXING Not everything is worth fixing. This phase is largely about saying which.
    Two of the five ranked items fall below the line, and the line is the point of the phase. A diagnosis that recommends everything has not been made, and the surfaces sitting in opposite quadrants get opposite treatments regardless of how similar their scores look.
  4. 04

    Design and build

    Prototypes and validation against the specific signals being targeted. Your team can build, or we can, or both.

    Produces: validated designs, acceptance criteria tied to signals. Takes: depends on scope.

    Phase 04

    Acceptance criteria tied to named signals

    TARGET SIGNAL Task completion, checkout SILENT FRICTION PASSES ONLY IF P50 attempts falls from 2.4 to under 1.5 P10 time to outcome falls below 90 seconds no new WCAG 2.2 AA failure introduced PROTOTYPE TEST · REVISE until the criteria are met, not until it looks finished A design with no named signal has no way to fail.
    Every design in this phase is aimed at a specific signal and carries numeric acceptance criteria written before the work starts. Without them, the review becomes a matter of taste and the only available verdict is whether the team likes it.
  5. 05

    Rescore

    Same model, same criteria, new evidence. The delta is the deliverable.

    Produces: rescored baseline, attributable delta per surface, updated confidence. Takes: one to two weeks.

    Phase 05

    The delta is the deliverable

    BASELINE · APRIL RESCORE · SEPTEMBER 2.8 @ 0.86 3.9 @ 0.88 +1.1 CHECKOUT +2.2 attributable SEARCH COPY +0.9 attributable BILLING no movement, treatment was wrong
    Same model, same criteria, new evidence. Because the criteria were fixed before the work, the movement is attributable per surface rather than claimed for the engagement as a whole, and a treatment that did nothing is reported as a treatment that did nothing.

What we need from you

  • Read access to product analytics
  • Access to support contacts or tickets from the last six months
  • Six to eight user interviews, which we can arrange or you can
  • Roughly four hours of one product person’s time across the engagement
  • Someone with the authority to say no to recommendations

That last one matters more than it looks. An engagement with no one authorised to decline produces a report nobody acts on.

Timeline and cost

A full first cycle runs six to ten weeks. A baseline plus diagnosis without the build phase runs three to four.

Tell us the shape of the problem and we will tell you honestly whether this is the right method for it. Or answer thirteen questions below and let the page tell you first.

The fit check

Find out whether we are any use to you

Thirteen questions about your product, your evidence and your decision-making. It returns one of twenty-seven results, and a good number of them tell you not to hire anyone, including us.

It is deliberately willing to talk you out of an engagement. If your sample is too thin, if the direction is already decided, or if the honest answer is that you should run the free framework yourself for a quarter, that is what it will say. Nothing you answer leaves your browser.