Observed behaviour enters at 1.00. Reported experience at 0.70. Anything inferred, including AI classification of tickets or session replays, at 0.50. Expert judgment, including a senior designer’s heuristic review, at 0.30. That last figure is the correct weight. Everyone in the field knows it, and no other framework says so numerically.
Confidence also decays, and shipping is what triggers it rather than the calendar. Each release touching a measured surface multiplies its confidence by 0.8, because evidence describes a product that no longer exists. The loop closes itself: shipping a fix lowers confidence, which pushes that surface back up the measurement queue, so re-measurement follows without anyone needing to remember it.