Operations

Measuring an analyst, not their output

Ask a transaction monitoring team how their analysts are performing and you will usually get a number: alerts closed per day. It is the easiest thing to measure and the least useful thing to know.

Throughput tells you how fast a queue is moving. It tells you nothing about whether the decisions in it were right, and it actively rewards the behaviour that makes them wrong. An analyst who closes forty alerts a day is either working a genuinely trivial queue or not looking properly, and the metric cannot distinguish between the two.

What you are actually trying to measure

The output of an investigation is a decision and a record of why it was made. Those are two different things and both can fail independently.

A decision can be correct with a record so thin that nobody can defend it eighteen months later. A record can be immaculate around a judgement that was plainly wrong. A quality framework that collapses these into one score will systematically miss one of them — usually the second, because good writing is persuasive.

So the first design principle: score the reasoning and the outcome separately, and weight them deliberately.

The categories that carry weight

Every QA framework I have built or reviewed eventually converges on a similar set. The labels vary; the substance does not.

  • Escalation judgement. Was the escalate/close call correct on the information available at the time? This is the heaviest single weight, and it must be assessed against what the analyst could see — not against what was discovered later.
  • Narrative quality. Does the write-up let an independent reader reconstruct the decision without asking questions? Not length. Reconstructability.
  • Evidence handling. Was the right evidence gathered, attached, and referenced? Missing attachments are the most common single defect and the easiest to fix.
  • Typology recognition. Did the analyst identify what they were actually looking at, or apply a generic close reason to a specific pattern?
  • Customer and context research. Was expected activity established before deciding whether activity was unexpected? Skipping this is the root cause of a surprising share of bad closes.
  • Procedural adherence. Codes, timeframes, referrals, the required fields. Low weight individually, but a persistent pattern here is a supervision signal.
  • Timeliness. Included, but weighted low and defined as within standard rather than fast. This is where throughput belongs — subordinate to quality, not alongside it.

Why weighting matters more than the categories

A flat framework where every category scores out of ten produces a number that is mathematically tidy and behaviourally useless. Analysts optimise for whatever is cheapest to improve, and the cheapest categories are always the procedural ones.

If escalation judgement is worth the same as filling in a close code, you will get excellent close codes.

Weight the framework so that the categories requiring actual thinking dominate the score, and so that a serious failure in judgement cannot be offset by tidiness elsewhere. In practice this means judgement and narrative together should account for roughly half the available marks, and no amount of procedural perfection should be able to rescue a case where the wrong call was made.

A scoring framework is not a measurement instrument. It is a statement of what the organisation rewards, and people will read it correctly.

The sampling problem nobody solves properly

Most QA programmes sample randomly, uniformly, at a fixed percentage. This is defensible and it is also inefficient.

Random sampling across a queue that is 90% low-complexity means 90% of your reviewer capacity is spent confirming that easy alerts were closed easily. Meanwhile the cases where judgement actually mattered — the borderline escalations, the unusual typologies, the high-value customers — get reviewed at the same low rate as everything else.

A better structure is stratified: a small random baseline across the whole population for defensibility and drift detection, plus deliberate oversampling of the segments where an error is expensive. Document the strata and the rationale. A supervisor will accept risk-based sampling readily; what they will not accept is sampling that looks risk-based but has no written basis.

Calibrate the reviewers before you trust the scores

The failure mode that quietly invalidates a whole programme: two reviewers score the same case differently, consistently, and nobody checks.

Once that is true, an analyst's score is partly a function of who reviewed them. Any decision made on those scores — coaching, ranking, performance management — is contaminated, and the analysts will work it out long before management does. Trust in the framework collapses, and it does not come back easily.

The fix is unglamorous and cheap. Periodically have every reviewer score the same set of cases blind, compare, and discuss the divergences. You are not looking for identical numbers; you are looking for divergences that reveal genuinely different interpretations of a category. Those get resolved in the guidance, not in the individual case.

What to do with the score

A QA score is a coaching input first, a control-health indicator second, and a performance-management input a distant third.

Used in that order, it works. Used in reverse, it degrades quickly: analysts become defensive about their reviews, reviewers soften scores to avoid conflict, and within two quarters you have a framework that produces high numbers and detects nothing.

The most valuable output is not the individual score at all. It is the pattern across the team. If six analysts independently mishandle the same typology, that is not six coaching conversations — it is a training gap or a procedure that does not say what people think it says. Aggregate defect analysis is where a QA programme repays its cost.

A minimum viable version

If you are starting from nothing, this is the smallest thing worth building:

  • Six to ten weighted categories, judgement-heavy, agreed in writing.
  • A fixed number of cases per analyst per month — consistency matters more than volume.
  • Stratified sampling with the strata written down.
  • A quarterly reviewer calibration exercise.
  • A monthly defect-theme summary that goes to whoever owns training.

That is a spreadsheet and a discipline, not a system. Firms spend a great deal of money buying the system before establishing the discipline, and the discipline is the part that produces the improvement.

More insights