Systems

Choosing a monitoring system: what to test before you buy

A vendor demonstration is a piece of theatre, and a competent one is very good. It runs on curated data, in a configured environment, presented by someone who has given it two hundred times. Everything works. That is not dishonesty — it is what a demo is for. It just tells you almost nothing about how the system will behave in your estate.

The questions below are the ones that produce differentiated answers.

Test with your own data, or do not test

The single most useful thing in a selection process is a proof of concept run on a real extract from your own environment — messy fields, missing values, your actual customer mix. Vendors will resist on effort grounds, and the ones who resist hardest are telling you something.

What you learn: how much data preparation the system genuinely requires, how it handles your gaps, and whether the alert volumes are anywhere near what was implied.

Ask who can change a rule, and how long it takes

This is the question that most affects your life for the next five years, and it rarely appears in a scoring matrix.

Can your team author and modify detection logic, or does every change require a vendor change request? What is the turnaround, and what does it cost? A system where a threshold change takes six weeks and a purchase order will quietly stop you tuning at all, and your calibration will rot.

You are not buying detection logic. You are buying how easily you can change detection logic once you learn what you actually need.

Look at the investigator's screen, not the dashboard

Selection committees are shown management dashboards. The people who will use the system for seven hours a day see the case screen.

Sit an actual analyst in front of it and have them work a case end to end. How many clicks to see the customer's history? Is the transaction context on the same screen as the decision? Can they attach evidence without leaving the case? Does the narrative field accommodate a real write-up?

Friction here compounds into thousands of hours a year and — more importantly — into worse decisions, because a tired analyst on a hostile interface takes shortcuts.

Interrogate the AI claims specifically

Every vendor now has one. Useful questions:

  • What exactly does the model do — rank, suppress, or detect something new? These are different risk profiles and vendors blur them.
  • What does it train on, and what happens in the first months when you have no disposition history in their system?
  • Show me a per-alert explanation as an investigator would see it, on a real case.
  • What documentation do you provide for model validation, and has it satisfied a regulator before?
  • If we turn the model off, does the system still work?

The last question is more revealing than it sounds.

Ask about exit before you ask about price

If you leave in five years, what do you get back? Rules in a readable format, or a proprietary export nobody else can ingest? Full alert and disposition history, or a summary? Who owns the tuning work you paid to produce?

Switching costs are the main reason firms stay on systems they have outgrown. The time to negotiate the exit is while they still want the deal.

Reference calls, asked properly

Vendor-supplied references are selected. Still worth taking, provided you ask questions that are hard to stage:

  • How long from contract to first production alert, and how did that compare to plan?
  • What did you discover after go-live that you wish you had tested?
  • How long does a rule change actually take?
  • What has the vendor been slow on?

Then find a non-supplied reference. Industry groups and the regional compliance community will tell you things a reference call will not.

The uncomfortable truth about fit

Most monitoring systems in the mid-market detect broadly similar things. The differences that matter over five years are configurability, investigator experience, data requirements and how the vendor behaves when something breaks.

Selection processes weight detection capability heavily because it is the easiest thing to score. It is rarely what determines whether the implementation succeeds.

Working through something similar in your own programme? I'm always happy to compare notes — get in touch.

More insights