Sanctions
Screening fuzziness is a policy decision — treat it like one
Somewhere in your sanctions screening configuration there is a number. It might be 85. It might be 82, or "medium". Whatever it is, it determines which name pairs the system considers close enough to stop and which it lets through.
Ask who set it and when, and the answer is usually some version of: it came with the system, and it has not been changed since implementation. Ask what risk appetite it represents and the question often does not parse, because nobody has framed it that way.
They should have. That number is one of the most consequential risk decisions the firm makes, and in most firms it was made by a vendor's default configuration.
What the threshold actually encodes
A fuzzy matching threshold is a statement about how much residual sanctions exposure the firm is prepared to carry in exchange for how much operational load.
Set it high — demand near-exact matches — and the alert queue is small, the false positive rate is tolerable, and some genuine matches with transliteration variance, reversed name order or a missing middle element pass through unstopped.
Set it low and you catch those, along with an enormous volume of noise that has to be worked by people. Below a certain point, the queue exceeds capacity, the average time per alert falls, and the quality of every decision in it degrades. A threshold set too low can produce worse outcomes than a moderate one, because a swamped analyst clearing on pattern is not really screening.
Both settings are defensible. Neither is defensible without evidence and a signature.
The question is never "is this threshold correct?" It is "can we show why we chose it, and what we accepted when we did?"
Why vendor defaults are the wrong answer
Not because they are badly chosen. Because they were chosen for a generic customer base that is not yours.
Matching performance depends heavily on the name population being screened. A book that is predominantly Anglophone personal names behaves entirely differently from one with significant Arabic, Chinese or Slavic transliteration. Names transliterated through different conventions produce variance that a threshold tuned on a different population handles badly in one direction or the other.
The same is true of entity names versus individual names, of markets where a small number of surnames are extremely common, and of any book with a meaningful volume of non-Latin-script origination.
A default is a reasonable starting point. It is not a calibration.
What a calibration exercise actually involves
The structure is the same regardless of platform. The work is in doing it honestly.
- Build a test set with known answers. Real names from your own population, plus deliberately constructed variants of listed names — transliteration differences, reversed order, dropped particles, common misspellings, punctuation variance. You need both true matches and true non-matches, and you need to know which is which before you start.
- Score the set across a range of thresholds. Not two candidate settings. A range, so you can see the shape of the curve rather than two points on it.
- Plot what you gain and what it costs. At each level: how many true matches are caught, how many are missed, and how many false positives are generated. The relationship is rarely linear and there is usually a region where a small reduction in threshold buys a large volume increase for almost no additional detection. That region is the finding.
- Test the variant categories separately. Aggregate performance hides the thing you most need to know — that the setting performs well on Latin-script names and poorly on transliterated ones, or vice versa.
- Decide, in a forum with authority. The output is a recommendation with quantified consequences, taken to whoever owns sanctions risk. Not an operational tweak made by whoever administers the platform.
Segmentation beats a single global number
The most common improvement available to a mature programme is not moving the threshold. It is having more than one.
A single global setting forces one compromise across populations with genuinely different characteristics. Screening a payment message against a list is a different problem from screening a customer record at onboarding: different data quality, different available context, different consequence of a miss.
Segmenting — by screening context, by list type, by name script or origin — lets each segment carry a setting appropriate to it. It is more work to justify and more work to maintain, and it usually produces both fewer false positives and better detection than a single number can.
The condition is documentation. Every segment needs a written basis. Segmentation with no documented rationale looks, on examination, exactly like someone quietly turning down the sensitivity where the queue was inconvenient.
What the file needs to contain
Assume you will be asked to justify the setting, by an examiner or by an internal audit function, roughly two years after the person who set it has left. The file needs to stand on its own:
- The test set, its construction method, and why it is representative of the book.
- Results across the tested range, not just the chosen point.
- The recommendation and the reasoning, including the volume consequences.
- Who approved it, in what forum, on what date.
- An explicit statement of the residual risk accepted — the categories of match that this setting is known to be weaker on.
- The review trigger: what would cause this to be revisited before the scheduled date.
That last item is the one most often missing and the easiest to specify. A material change in customer geography, a new list with different naming conventions, a platform upgrade that changes the matching algorithm, a spike in confirmed near-misses — each should force a re-look regardless of when the last one was.
The honest limitation
No threshold catches everything, and a calibration exercise does not change that. What it changes is whether the misses are known and accepted or unknown and discovered by someone else.
That distinction is most of what separates a finding from a serious finding. A firm that can produce a calibration file, state which match categories its setting is weak on, and show that it made that trade deliberately at the right level of seniority is in a fundamentally different position from one whose answer is that the number came with the system.