PONOPT FIELD NOTES · HR и workforce operations

Workforce Metrics That Improve Service Instead of Encouraging Gaming

Speed, CSAT, and QA scores quietly reward gaming when promoted alone. Learn paired, outcome-based metric design that actually improves frontline service.

The fix for gaming is not fewer metrics but better-designed ones. Average handle time, CSAT, and internal QA scores each reward the cheapest shortcut when promoted alone: fast replies that defer work, surveys sent only to easy cases, and inflated grading. Service improves when you pair every efficiency metric with an outcome counter-metric, measure resolution through observed repeat contacts, calibrate evaluators, fix survey hygiene, and treat contradictions as incidents rather than wins.

Key takeaways

  • Any measure turned into a target stops being a good measure (Goodhart's law); employees optimize the number rather than service, and that is rational adaptation, not malice.
  • AHT, CSAT, and QA scores are the most common anti-patterns when pushed in isolation, producing fast touches, premature closure, survey sampling shifts, and grade inflation.
  • The warning signal is not a falling number but a contradiction: a primary metric improving while two outcome indicators worsen for two straight weeks.
  • Before scaling any improving KPI, translate metric → behavior → customer impact and classify the cause as gaming, measurement drift, or real improvement with lag.
  • Reliable resolution is measured by observed repeat contacts in a 24–72 hour window, not agent self-report; QA grading needs multi-evaluator calibration.
  • A counter-metric without an owner and a pre-decided threshold is just trivia; define in advance what pauses when the outcome metric worsens.

The anti-pattern: when the measure becomes the goal

Economist Charles Goodhart observed in 1975 that any statistical regularity tends to collapse once pressure is placed on it for control purposes; anthropologist Marilyn Strathern later condensed the idea into the now-common phrasing that once a measure becomes a target, it ceases to be a good measure. This is not an abstraction about monetary policy but an everyday fact of frontline operations: the moment a leader announces a drive to cut average handle time, agent behavior begins to change before service does.

It helps to stop reading this as an accusation against staff. When bonuses, ratings, or even job security are on the line, people optimize what is rewarded as naturally as water finds a crack in pavement. The first rule of metric design is therefore to make good service cheap and gaming expensive. If the easiest path to the target harms the customer, the target is badly specified — the problem is the incentive, not the person.

  • Measurement changes behavior before it changes outcomes — design with that in mind.
  • Gaming is usually a rational response to an incentive; fix the incentive, not the people.
  • The cheapest route to any number almost always exists — and it almost always costs the customer.

Common metric anti-patterns in frontline service

Average handle time as a headline target is the classic. Under speed pressure, agents learn to make short touches: a macro acknowledgement instead of diagnosis, a fast reply that advances the ticket an inch. AHT falls while repeat contacts within a week and escalations climb; total effort for the customer and the team rises even though the slide shows a win.

CSAT distorts through sampling. If surveys fire only on closed, "resolved" tickets or you change timing and eligibility, the average can rise while the response rate falls and the free-text comments get angrier. That pattern is usually a reporting problem, not proof that customers are complicated.

Internal quality scores suffer from grade inflation: reviewers pick easy tickets, mark leniently to avoid harsh binary judgments, and agents learn to "write to the rubric" — hitting tone and formatting while missing the customer's actual goal. The result is a confident-looking 90%+ score that does not track real experience. Remote work adds activity metrics such as hours on screen and keystrokes, which employees game with performative work because an unfair system does not feel worth respecting.

  • AHT alone → fast touches, premature closure, rising repeat contacts and escalations.
  • CSAT without sampling hygiene → rising score with falling response rate and angry verbatims.
  • QA scores without calibration → inflation, easy-ticket grading, and playing to the rubric.
  • Activity metrics (hours, keystrokes) → performative work instead of results.

Recognizing a metric mirage within 48 hours

The key signal is not a bad trend but a contradiction: a primary metric improving while two outcome indicators worsen for two consecutive weeks. Build a contradiction snapshot: the improving KPI, one adjacent operational metric, and three downstream indicators — escalation rate, repeat contact within seven days, reopen rate, refunds, or complaint volume. Add a few artifacts: several escalated threads, one QA review, and one customer verbatim.

Then classify the cause instead of blaming anyone. Three problems look identical on a dashboard. Gaming: the system finds the cheapest path by pushing work downstream or off-metric. Measurement drift: the definition, sampling, eligibility, or channel mix changed rather than the experience. Lag: the KPI leads and the outcome trails, so today's improvement may take weeks to show up in churn and renewals. Run one disconfirming test per hypothesis: sample fast-closed tickets for empty macros and thin notes, hold cohorts stable across channels and ticket types, then give nearer-term outcomes one to two weeks to move before churn does.

The decision rule: do not change targets broadly until you can say in one sentence which cohort got worse and what behavior likely shifted. This prevents both ripping out a change that is genuinely working and institutionalizing a harmful "this is how we do support now" habit.

Designing metrics that do not invite gaming

The answer is not fewer metrics but linked ones. Tie speed to resolution quality: rather than merely adding a reopen chart, introduce a closure review at the exact pressure point where tickets get closed to protect AHT or SLA. Every counter-metric needs an owner and a pre-decided threshold — what gets paused, rolled back, or reviewed when the outcome metric crosses the line. Without that, a counter-metric is just a decorated dashboard.

Measure resolution through observation, not self-report. Reliable first-contact resolution is the share of contacts after which the customer does not return with the same issue within a defined window — typically 24 to 72 hours — which requires linking contacts across channels. Agent self-reports of "resolved" are gameable under performance pressure and should not be treated as evidence.

Outcome metrics — repeat contacts, reopens, refunds, churn — are harder to game but lag behind. Use them as the ultimate scoreboard for success, and fast operational metrics as the steering wheel for daily management. The pairing is the protection: gaming one metric becomes expensive because it damages the other.

  • Paired metrics: speed ↔ repeat contact, deflection ↔ escalation, SLA ↔ resolution quality.
  • Measure FCR by repeat contact in a 24–72 hour window, not by agent self-report.
  • Outcome metrics (churn, refunds) are the scoreboard; operational metrics are the daily steering wheel.
  • Every counter-metric needs an owner and a pre-defined trigger threshold.

Measurement hygiene: calibration and stability

Quality assurance is only as good as its scoring process. Run regular calibration sessions where several evaluators blindly score the same interactions and reconcile their differences; without this, supervisors drift apart over months and agents receive different ratings for identical work. Let evaluators quote the transcript, which reduces the grade inflation that comes from scoring from memory rather than evidence.

Pin the survey send rules for the quarter. If you change CSAT timing, eligibility, or sample, annotate the dashboard and treat the trend as discontinuous until you rebuild the baseline. Define "repeat contact" — a customer returning within seven days about the same issue — and keep the definition stable rather than adjusting it to flatter reporting. Changing the definition changes the number, not the service.

Before attaching any metric to compensation, run the translation metric → behavior → customer impact: what would an agent under pressure do differently today, and which harmful shortcut appears first? Once a metric is tied to pay, people optimize it as if their rent depends on it, because it does. Design guardrails before that point, not after.

Practical decisions and trade-offs for leaders

Adopt a speed goal only alongside a definition of "resolved outcome" for your top ticket types, and coach to that outcome rather than to "short reply." Speed should come from better diagnostics, macros, and routing, not from skipped steps. Where a fast first response matters, require the first reply to include one concrete next step that moves the issue forward, not an empty acknowledgement.

Tighten closure policy surgically: require a closure reason that matches an outcome and add a reopen-review loop oriented toward learning rather than punishment. Treat deflection as containment with resolution, not simply fewer tickets: track escalations from deflected sessions and complaints surfacing in other channels.

Do not scale a local win into a compensation program without an audit. A local improvement becomes a company narrative, then a comp plan, then a permanent source of strange behavior. Use stable cohorts and A/B tests before rolling targets or pay structures out broadly, and reserve judgment until lagging outcomes confirm the win.

  • Speed only with a "resolved outcome" definition and a concrete next step in first replies.
  • Closure by outcome-based reason, with a learning loop for reopens.
  • Deflection as containment with resolution, controlling for escalations.
  • Do not scale a "win" into pay without an audit and cohort checks.

Metric design guardrail checklist: audit before a KPI reaches pay

A fast, repeatable audit for any primary KPI you intend to promote, tighten, or attach to compensation. Work through the items in order; if you cannot answer one, the target is not ready to scale.

  1. State in one sentence: "If I am an agent under pressure trying to look good on this metric, what do I do more of, and what do I do less of?"
  2. Name the cheapest path to the number, including the one that harms customers (fast close, empty macro, hidden contact).
  3. Add a quality counter-metric and a downstream outcome indicator (7-day repeat, escalation, refund).
  4. Freeze the definition, sampling, and send rules for the quarter.
  5. Define the counter-metric's trigger threshold and what gets paused when it is crossed.
  6. Assign an owner and a review trigger (new policy, routing change, tightened target).
  7. Check FCR via repeat contact in a 24–72 hour window, not agent self-report.
  8. Run QA calibration: one interaction scored blindly by several evaluators, then reconcile variance.
  9. Compare like-for-like cohorts (channel, type, tier) before and after any change to exclude drift.
  10. State the expected lag of the outcome metric (churn, renewals) in weeks, not days.

Questions people ask

Why does average handle time encourage gaming?

When AHT becomes the single goal, agents get an incentive to make short touches instead of fully resolving issues: macro acknowledgements, deferring work to a later contact, or premature closure. AHT falls while repeat contacts within a week, escalations, and total customer effort climb. This is Goodhart's law in action: the number improves because the workflow learned to produce it, not because the customer's outcome improved. The fix is not to abandon speed but to pair it with a resolution-quality counter-metric and require a concrete next step in every first reply.

How can I stop CSAT from being gamed through sampling?

CSAT is easy to inflate by changing who gets surveyed and when, such as firing surveys only on solved tickets or after successful outcomes. The tell is a score that rises while the response rate falls and free-text comments turn more negative. Fix survey hygiene: freeze the send rules for the quarter, do not change eligibility or timing without annotating the dashboard, watch response rate alongside the average, and read verbatims rather than only the mean. If you change rules, treat the trend as discontinuous until a new baseline is rebuilt.

What is QA calibration, and why should scores above 90% look suspicious?

Scores above 90% often reflect a flawed QA program rather than excellent service: graders favoring easy tickets, lenient marking to avoid harsh binary judgments, and agents learning to write to the rubric while missing the customer's goal. Calibration is the process of having several evaluators blindly score the same interaction and then reconcile their differences, aligning their understanding of criteria. Without regular calibration, evaluators drift apart over months and agents receive different ratings for identical work. Calibration is the main defense against grade inflation and inconsistent scoring.

Why is deflection considered a vanity metric, and how do I protect it?

Deflection measures activity (how many contacts were pushed away) rather than the outcome. If you deflect by hiding human contact options or aggressively blocking access, demand does not vanish — it mutates into escalations, refunds, and complaints in social channels. Treat deflection as containment with resolution: track escalation rates from deflected sessions, repeat contacts, and feedback in other channels. High deflection alongside rising complaints is a classic sign you are optimizing the wrong metric, not improving self-service.

How should first-contact resolution (FCR) be measured reliably?

Reliable FCR is the share of contacts after which the customer does not return with the same issue within a defined window, usually 24 to 72 hours, and it requires linking contacts across channels to detect a repeat about the same topic. Agent self-reports of "resolved" are gameable under performance pressure and unreliable. Segment FCR by channel and agent cohort, since voice typically outperforms asynchronous channels and results differ between new hires and veterans, revealing exactly where coaching or workflow fixes belong.

How is a counter-metric different from an extra dashboard chart?

A counter-metric makes gaming the primary metric expensive by penalizing a linked outcome: speed paired with repeat contacts, deflection with escalations, SLA with resolution quality. But without an owner, a pre-decided threshold, and a rule for what pauses when the counter-metric worsens, it is just trivia with better branding. Decide in advance which incentive you will freeze or roll back when the outcome metric crosses the line, assign a responsible owner, and review it on triggers rather than only on a calendar.

Sources and further reading

Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.

  1. When Metrics Improve but Outcomes Get Worse: Diagnosing Goodhart Problems Before They SpreadCalypso
  2. How Teams Accidentally Optimize the Wrong Metric (and How to Catch It Early)Calypso
  3. Remote work brought unfair performance metrics – and employees are gaming themMcGill University
  4. Dangers of the 90%+ QA ScoresMaestroQA
  5. What Is Goodhart's Law? Balancing Authenticity & MeasurementBMC
  6. First Contact Resolution (FCR) — GlossaryKustomer
  7. Gaming the System: What a USPS Smiley Face Reveals About Bad MetricsLean Blog (Mark Graban)