When the user cannot leave: measuring internal tools

In three points

  • Every conventional UX signal is a proxy for exit. An internal tool removes the exit, so the whole instrument panel reads clean while the cost keeps accruing somewhere it is not counted.
  • HR systems are the extreme case: one interface serving a daily operator, a quarterly manager and an employee who meets it five times a year, with only the first of the three in the research.
  • Perception inside an employer's system is collected under a power asymmetry, so it is reported evidence that needs a further discount rather than a flat 0.70.

Churn is flat, because leaving is not on offer. Abandonment is near zero, because the form is compulsory. Ticket volume is low, because people ask the person two desks away. The annual engagement survey rates the HR system 4.1 out of 5. Every instrument you own reports a healthy product, and the finance team is still paying for a process that takes four minutes when it should take forty seconds.

This is not a measurement failure at the margins. It is the whole instrument panel failing at once, for one reason: the signals were designed for markets, and an internal tool does not have one.

Fig. 01

The signals all assume an alternative

EVERY UX SIGNAL IS A PROXY FOR EXIT NONE APPLY CHURN NOT AVAILABLE the only way out is to leave the company ABANDONMENT NOT AVAILABLE the form is mandatory, so the user returns TICKET VOLUME UNLOGGED they ask a person, not a queue SATISFACTION BIASED UPWARD the employer is the one asking WHERE THE COST GOES INSTEAD SHADOW SYSTEMS DEADLINE CLUSTER ASSISTED ENTRY QUIET ATTRITION
A consumer product leaks users, which is loud and countable. An internal tool leaks hours, accuracy and eventually people, none of which arrive in a dashboard with the product's name on it. The four boxes are where to go looking, because the cost is conserved even when the signal is not.

Every signal you own is a proxy for exit

Churn measures a decision to stop paying. Abandonment measures a decision to stop trying. Ticket volume measures a decision that the product was insufficient and a stranger might help. Satisfaction measures a willingness to say so. All four are downstream of the same thing: the user had somewhere else to be, and the option to go there is what converts their experience into a number you can read.

Take the option away and the numbers do not degrade gracefully. They invert. Completion goes up in a bad internal tool, because the employee who cannot file the expense keeps trying until it works. Retention is total. Ticket volume falls as the tool gets worse, because people learn that the queue is slower than the colleague. A dashboard built on those four signals will show an internal system improving as its users get more resigned.

Where the cost goes instead

Cost is conserved. If the user cannot spend it on leaving, they spend it on four other things, and each leaves a trace somewhere outside the product analytics.

Shadow systems. The spreadsheet running alongside the system of record. Every internal tool of any age has one, maintained by someone who will explain it as a personal preference. It is not a preference, it is a bill. Find it by asking a team to walk you through the process rather than through the software.

Deadline clustering. When a task is compulsory and unpleasant, it is deferred to the last legal moment. The distribution of submission timestamps against the deadline is a friction reading nobody thinks to take: a task people find cheap is spread across the window, a task people dread arrives in the final few hours as a spike.

Assisted entry. The share of records created by somebody other than the person they describe. When an HR operator files a request on an employee's behalf, the product has failed and the failure has been absorbed by payroll rather than by support. This is the single most underused number in internal tooling, and it is usually already in the database.

Quiet attrition. The slowest and most expensive of the four, and the only one that eventually shows up in a business review. By the time it does, it has been attributed to compensation.

Why HR systems are the extreme case

Most internal tools serve one population doing one job. An HR system serves three, whose exposure to the interface differs by two orders of magnitude, and it serves them through the same screens.

The HR operator lives in it. Several hundred sessions a year, complete mastery, a personal set of workarounds, and a well-formed opinion. The manager arrives quarterly, under a deadline, to approve things. The employee arrives about five times a year: a time-off request, a document to sign, an address change, an annual review they have been reminded about four times.

The employee never amortises the learning curve. Every visit is a first visit, because the twelve weeks since the last one were enough to forget the layout, the vocabulary and where the button was. Design decisions that are efficient at four hundred repetitions, such as dense tables, terse labels and required internal codes, are hostile at five. The operator experiences the system as a tool. The employee experiences it as a test.

Fig. 02

Exposure inverted against who was asked

ONE SYSTEM, THREE POPULATIONS ONE INTERFACE HR OPERATOR~1,100 / yr learns it, then builds workarounds around it MANAGER~14 / yr approves under deadline, never learns it EMPLOYEE~5 / yr meets it new almost every time EXPOSURE IS NOT WHAT DECIDES WHOSE VIEW IS COLLECTED POPULATION RESEARCH OPERATORS 4% 92% EMPLOYEES 96% 8%
The operator is articulate, available, and the person the vendor or the internal team already talks to. That is why the research lands there. It is a sample drawn from four per cent of the population, describing an interface the other ninety-six per cent experience completely differently.

The perception reading needs a second discount

RUCF weights reported evidence at 0.70 because people misremember, rationalise and round. Inside an employer's own system, a third mechanism is operating: the answer has consequences.

An employee asked to rate the performance review tool is answering a question about performance review, from an account with their name on it, in a survey administered by the function that runs their file. The instrument is not neutral and neither is the respondent. Complaints about an HR system are heard, correctly or not, as complaints about HR, and the people who have learned that stop making them. This is the same mechanism that produces silent friction in consumer products, with a power relationship added on top of it.

So treat internal perception as reported evidence at 0.70 with a further multiplier for collection conditions: 1.0 where a third party collects it and the responses are genuinely anonymous, 0.8 where the employer collects it anonymously, 0.6 where responses are attributable. Declare the multiplier next to the score. If that produces a perception reading too weak to diagnose against, the finding is that you do not currently have a perception channel, which is worth knowing before you act on the one you thought you had.

The behavioural side is unaffected, and this is the compensation. Internal systems are usually instrumented far better than consumer products, because they run on infrastructure you own, against identities you control, for populations you can enumerate exactly. The friction measure can be very strong here even when the perception measure is very weak.

The four currencies, translated

The four currencies still apply, but two of them need restating for a population that cannot walk away.

CurrencyInternal formWhere it lives
Time to outcomeIntent formed to record filed, including the deferral windowObserved
AttemptsDistinct formulations, plus corrections made after submissionObserved
AssistanceThe HR inbox, the team channel, the colleague who knowsReported floor
AbandonmentDeadline clustering and assisted entry, since leaving is not permittedObserved

Time to outcome is the one most often mismeasured, because the clock is started at the page load. The real window opens when the employee decides to request the leave and closes when the request is filed, and in a badly designed flow most of that window is spent not using the product at all: waiting to be told the cost centre code, finding the policy document, asking whether this counts as sick leave. That waiting is the product's cost even though the product was closed for all of it.

Why the internal-tool weights look like that

The scoring model weights an internal tool Task 0.40, Access 0.30, Commitment 0.15, Belief 0.15. Both of the outliers are deliberate.

Task at 0.40 is the highest of any product type, because filing the thing is the entire reason the system was bought. Nobody opens an HR platform for pleasure, browses it, or recommends it. There is one job and either the interface does it or a person does it instead.

Access at 0.30 is also the highest of any product type, and not because internal populations skew disabled. It is because they are captive. A disabled user of a public product has competitors; a disabled employee has an employer, and a keyboard trap in the approval queue is a condition of their employment rather than an inconvenience. Conformance remains a gate rather than a score everywhere in RUCF, and in an internal tool the gate is also a legal exposure with a named claimant.

Commitment at 0.15 is low because return is compulsory, so retention measures the employment contract rather than the product. It is not zero, because depth of use still detects the shadow system: a population using one screen out of eleven is telling you exactly where the other ten went.

Belief at 0.15 is low because legitimacy is conferred by the employer rather than earned by the interface. It stays in the model because distrust in an internal system is expensive in a specific way: employees who do not believe the payroll record duplicate it, and now you have two records and no way to know which is wrong.

Two things this cannot do

It cannot price the finding in revenue. There is no conversion event to multiply against, so an internal-tool cost is priced in hours and in headcount: minutes per submission times submissions per year times loaded hourly rate, plus assisted entries times operator time. That arithmetic is less impressive than a revenue number and considerably harder to argue with, since every term in it is already in a system somebody owns.

And it cannot tell you whether the process should exist. A system scoring 4.4 that efficiently administers an approval step nobody has justified since 2019 is a well-measured waste. RUCF measures what the interaction costs the people who have to have it. Whether they should have to have it at all is a different question, asked of a different document.

Signals this affects

Primarily the Task surface: task completion, error recovery and navigation clarity, with the Access surface carrying unusual weight. The collection-conditions discount modifies evidence confidence in Layer 2.

Find your gap


Related: Case study: an HR platform whose buyer loved it · Silent friction: the failures no survey will find · Accessibility is a gate, not a score