The customer-service team had a problem.
Tickets were taking too long to close. Some had been sitting in the queue long enough to qualify for a pension.
So management introduced a simple measure: average resolution time.
This was sensible. The number exposed neglected tickets, made teams comparable and gave everyone a problem they could actually see. Within months, resolution time fell dramatically.
Customers were still unhappy.
They had simply learned to open a second ticket after the first one was closed without solving the problem.
Nobody sent an email saying, “Please become less helpful.” The organization merely made speed visible, comparable and professionally consequential.
The people did the rest.
We need measurements
It is easy to make fun of metrics after they go wrong.
It is harder to run a school, hospital, company or government without them.
Intuition does not scale particularly well. The manager who “just knows” who is productive may be measuring confidence, similarity or who sends messages at 11:43 p.m. A number can reveal failure that hierarchy would prefer not to see.
The problem begins when a useful description becomes a target.
This family of problems is usually gathered under Goodhart’s law: once a measure becomes a target, it tends to lose some of its value as a measure. A related idea, Campbell’s law, warns that the more a quantitative indicator is used for social decision-making, the more pressure there is to corrupt both the indicator and the process it monitors.
The familiar explanation is gaming.
People manipulate the numbers.
That happens. But I think it is the less interesting half of the story.
A metric has a career
Most troubled metrics begin innocently.
Working model—not measured data
indicator→Target and
consequences→Attention
redirects→Unmeasured work
loses ground→The proxy
becomes the job
1. The metric points at something useful
Ticket time roughly indicates responsiveness.
Test scores contain information about learning.
Hospital waiting times matter to patients.
Story points help a software team discuss the size of work.
None of these measures is meaningless. If it were, nobody would adopt it.
2. Consequences attach themselves
The number appears in the monthly review.
Then on the executive dashboard.
Then in performance evaluations.
Soon it affects budgets, bonuses, public rankings or whether someone has to explain a red box to a vice-president.
The metric has stopped being a thermometer. It has become a thermostat.
3. Attention moves toward the measurable surface
Teachers spend more time on tested material.
Developers become remarkably philosophical about whether a task is three points or five.
Hospitals reorganize activity around a waiting-time threshold.
Employees produce progress that photographs well for the status meeting.
Some of this is deliberate. Much of it is simply adaptation.
Research on public-sector performance systems has documented patterns with wonderfully gloomy names: tunnel vision, measure fixation, myopia, sub-optimization and gaming. The point is not that public servants are unusually devious. It is that a narrow signal becomes the shared map for a complicated job. (Propper & Wilson, 2003; Bevan & Hood, 2006)
4. Unmeasured work becomes harder to defend
Imagine two support agents.
One closes twelve straightforward tickets.
The other spends the afternoon tracing a strange failure that will eventually prevent hundreds of tickets.
The dashboard knows exactly what the first person did.
The second person has a story.
Stories do poorly against columns.
Over time, the organization may not merely reward visible work. It may lose the language required to value invisible work.
5. The proxy becomes the job
At this point, people no longer experience the metric as an approximation.
The ticket must close.
The points must increase.
The target must be met.
The metric originally measured reality.
Then it reorganized reality.
Honest people can still distort a system
Suppose a school places heavy weight on a standardized test.
A teacher does not need to cheat for the curriculum to narrow. She knows which skills will be tested, which students are close to a threshold and which activities will not appear in the results. Faced with limited time, she redirects effort.
Every choice may be locally defensible.
The aggregate effect may still be a school that has become better at producing the evidence of learning than learning itself.
The same thing happens in less dramatic ways at work. Once a number controls meetings, praise and promotion, it becomes an attention-allocation device. It tells people which part of reality is worth noticing.
Eventually, they may sincerely believe that part is the whole.
This is why “just hire ethical people” is not a complete metric strategy. Character matters. So does the environment that repeatedly teaches character what counts.
More metrics are not automatically the cure
The obvious solution is a balanced scorecard.
If speed distorts behaviour, add customer satisfaction. If satisfaction is gamed, add repeat contacts. Then quality, cost, employee wellbeing and a seventeen-tab workbook nobody can open without enabling macros.
Multiple measures can help because they make it harder to optimize one narrow surface.
They can also create new problems:
- people choose the easiest metric in the bundle;
- leaders combine unlike measures into one suspiciously precise score;
- documentation expands while judgment shrinks;
- and nobody remembers which outcome the dashboard was supposed to protect.
A better defence may involve several different kinds of evidence:
- a small set of quantitative measures;
- qualitative review of actual cases;
- measures of likely side effects;
- periodic rotation so one proxy does not become the permanent definition of success;
- and enough professional discretion to say, “The number improved, but the thing got worse.”
Most importantly, the organization can periodically test whether the proxy still predicts the outcome.
Do faster ticket closures still produce fewer unresolved problems?
Do higher scores predict later understanding?
Do more releases produce more user value?
A metric should have to reapply for its job.
Sometimes the number really helps
This argument can overcorrect.
Targets have improved performance in some settings. Waiting-time measures can expose neglect. Safety checklists can make critical omissions visible. Sales numbers can tell us whether anything was sold, which remains an underappreciated feature of a sales department.
The danger is not measurement.
It is attaching strong consequences to a narrow proxy while treating adaptation as a character defect rather than a predictable system response.
Maybe every metric contains a bargain:
I will simplify reality enough to help you coordinate, if you promise not to confuse me with reality.
Organizations are very good at accepting the first half.
I am less sure about the promise.
What I still cannot figure out
Rotating metrics, combining evidence and preserving judgment can slow the corruption.
But people are excellent learners. Given enough time, they discover what the institution celebrates—even when the institution never states it plainly.
Can an organization use a metric long enough to improve performance without eventually teaching people how to perform the metric?
Until the next strange question,
Osagie
My current hunch: A metric becomes dangerous when it stops informing judgment and starts replacing it.
Most likely reason I am wrong: Some outcomes are measurable enough, and some feedback loops fast enough, that a strong target remains a better guide than human discretion.