Perspectives
Perspective 018Technology

The biggest AML risk with AI is not hallucination. It is plausibility.

A weak analyst rationale often looks weak. AI can make weak reasoning look convincing. A practitioner's view on what an MLRO should see before approving an AI-enabled AML control for production.

By Everett MorganAutumn 202612 min read
An empty chair at a boardroom table beside a switched-off screen at dusk
Executive brief

Reading time · 15 minutes  ·  Primary audience · Boards, MLROs, Heads of Financial Crime, Compliance Officers, Chief Risk Officers

Why this matters

AI has legitimate uses in AML work. There is no virtue in asking a skilled analyst to spend an hour assembling information that technology could assemble in seconds. But removing manual effort and transferring judgement are two very different things. The closer AI moves towards influencing customer risk, alert disposition, escalation or suspicious-activity decisions, the more I want to see the evidence underneath it. Not the vendor presentation. Not the policy. The evidence.

Key findings
  1. 01Plausibility is a more difficult control problem than obvious error. AI can improve the presentation of weak reasoning without improving the reasoning. A fluent rationale is not necessarily a defensible one.
  2. 02Governance should be proportionate. The depth of validation, evidence and oversight should follow the materiality of the use case and how directly it influences AML decisions. Proportionality is not a basis for having no governance.
  3. 03Perfect technical interpretability may not be achievable. Decision traceability and governance accountability still are, and that is where the firm's obligation sits.
  4. 04You cannot govern AI use you have not identified. A formal AI governance framework does nothing about the analyst who is quietly using an external tool. Shadow AI has to be found first.
  5. 05Human in the loop is not enough. It tells us where the human sits. It does not tell us what the human contributes, and no override rate is inherently good or bad.
  6. 06Model approval is not permanent, and AML has no clean ground truth. Drift has to be read from several indicators together rather than one metric.
  7. 07The objective is not zero AI risk. It is understood AI risk, consciously accepted by the people accountable for the control.
Questions Boards should ask
  1. 01Which material financial-crime decisions does AI currently influence, and where does final human accountability sit?
  2. 02What evidence tells us that the AI-enabled control is performing within the limits we approved, and what would cause us to restrict or stop it?
  3. 03If the vendor and the model were unavailable tomorrow, could we reconstruct a material AI-influenced decision from our own records?
  4. 04What human review has been removed to create the efficiency benefit, and what residual risk have we accepted in return?

A convincing wrong answer is harder to catch

I am less worried about an AI tool producing an obviously ridiculous answer than I am about it producing a very convincing wrong one.

In an AML function, that distinction matters.

A poor analyst rationale can usually be spotted. The evidence is thin. The customer profile has been misunderstood. The transaction analysis does not support the conclusion the analyst reached. Someone reviewing the file can see that more work is needed, and can say so.

AI changes that problem. It can take weak analysis and give it structure, confidence and reasonable-sounding language. The reviewer then has to answer a harder question. Is the reasoning good, or does it simply read well?

I should be fair about the title. Plausibility is not the only significant AI risk in an AML function, and it will not be the dominant one in every deployment. Data quality, bias, integration risk, model risk, cyber and operational resilience, third-party dependency and regulatory risk are all real, and in some deployments one of them will matter more.

Plausibility earns particular attention for a different reason. It is hard to detect. The output can look strongest precisely when the reasoning underneath it is weakest, and the usual review controls are built to catch work that looks thin.

Why polished output receives less challenge

This is where the human factors matter, and they are usually left out of the discussion. A plausible AI rationale becomes considerably more dangerous when it lands on an analyst who is carrying too many cases, working to a productivity target, newly trained or not trained at all, using a workflow that makes accepting the output easier than questioning it, or working in a firm that has been told repeatedly how good the technology is.

Automation bias is not a theoretical concept. A rationale that arrives polished, structured and complete may receive less challenge than a rough one, precisely because it looks authoritative. That is the opposite of what the control was designed to do.

That is the question this Perspective is about. It is also the first of a series looking at AI in AML and financial crime work.

AI assistance is not AI decision-making

Two instructions that look similar and are not. Summarise this customer's activity. Does this activity look suspicious?

Between those two there is a progression, and firms move along it faster than their governance does. Information gathering. Summarisation. Prioritisation. Recommendation. Drafting the rationale. Risk scoring. Alert disposition. Escalation. Customer restriction. Support for the SAR or STR decision.

That is not a regulatory taxonomy and I am not offering it as one. The point is simpler. The closer the technology moves to judgement, the more evidence I want before I approve it.

And I do not accept: “The system decided.” That is not an AML governance answer.

Proportionality, before anyone builds a committee

Everything I set out in this Perspective should be read through proportionality. I am not suggesting the same governance architecture for every firm or every use case, and I would resist anyone who read it that way.

The depth of validation, evidence, monitoring and oversight should follow the risk. In practice I look at the materiality of the use case, how directly the technology influences AML decisions, the financial-crime risk in the underlying activity, the potential customer and regulatory impact, the complexity of the technology, the size, nature and operational complexity of the firm, and whether other controls already cover the same exposure effectively.

A Tier 1 bank placing AI inside transaction-monitoring decisioning should not have the same governance as a smaller regulated firm using a tool to tidy up documentation or summarise a low-risk administrative task. The first needs pre-deployment testing, case-level traceability and formal performance monitoring. The second may need a defined purpose, a data boundary, an owner and somebody sensible checking it occasionally.

One caveat, because I have watched this word get misused. Proportionality is a basis for scaling governance. It is not a basis for having none.

Before production means before production

When a deployment is put in front of me, these are the questions I ask, in roughly this order.

None of that requires a data science background. It requires somebody to insist on answers before the tool starts touching customers.

The AML risk assessment should not become the document written six months later to explain why the model behaved differently from what everybody expected. The risk assessment belongs before production, not after the first control failure.

Classification under the AI Act depends on the actual intended use. Not every AI use in a financial crime function falls into the same category, and I would be wary of anyone who tells you that all AML AI is automatically high-risk, or that a model you cannot fully open up is automatically unlawful. What the tool is designed to do, and what it is actually being used for, drive the answer.

One point of discipline while we are here. Three things get blurred together in AI governance papers and they should not be. There is what the law actually requires, in AMLR, the AI Act, DORA and GDPR. There is what supervisors expect or set out as principle, including EBA governance, outsourcing and ICT material. And there is what I would recommend as a practitioner. Most of the questions in this Perspective sit in the third category. They are not statutory obligations and I would not present them as if they were.

The AI I worry about first may not be the one Procurement bought

The analyst who puts a customer profile into an external chatbot because it writes a cleaner summary. The relationship manager using an unapproved tool to rewrite a source-of-wealth analysis. The investigator pasting transaction details into an external model to see what it makes of them.

Then the questions start.

I would not label every instance a data-protection breach. That depends on the facts. But it is potentially a material data-protection, confidentiality, information security and AML governance risk at the same time, and most firms discover it late.

There is a second problem underneath the first. If analysts are using unapproved tools to help with CDD, alerts, investigations or SAR and STR analysis, the firm does not know which AML decisions have been AI-assisted. It cannot describe its own control.

You cannot govern AI use you have not identified.

The Empty Chair Test

To be clear about what this is not. It does not mean the firm needs the vendor's source code, proprietary model architecture or intellectual property. I have never asked for that and I would not expect to get it. The objective is not complete technical independence from the vendor.

The objective is avoiding a situation where the firm cannot explain or evidence its own AML control because the explanation only exists inside somebody else's company.

For a material AI-influenced AML decision, I would expect the firm to be able to retrieve, where relevant to the case and the design of the tool:

I am not asking for SQL access, logits, feature vectors, SHAP values or one particular technical architecture, and I am not asking anyone to reverse engineer the model. The requirement is technology-neutral. The firm needs enough case-level evidence to reconstruct the decision path.

If we cannot retrieve that record without asking the vendor to reconstruct it after the event, I do not think the control is sufficiently auditable. At that point we have not simply bought technology. We have outsourced part of our ability to explain our own control.

Human in the loop is not a control metric

Human in the loop tells me where the human sits. It does not tell me what the human contributes.

What tells me something is the reviewer behaviour: acceptance, amendment, override, escalation, how long the review took, patterns by analyst and team, differences across risk cohorts and typologies, and the quality of the rationale that came out the other end.

There is no universally correct acceptance, amendment or override rate, and I would not trust anyone who offers one. A 5% override rate could mean the model performs exceptionally well. It could equally mean the reviewers have stopped looking. A 60% override rate could mean the model is poor. It could also mean human challenge is working exactly as intended.

So the governance question is not whether a percentage is inherently good or bad. It is whether the firm understands the pattern, investigates material changes in it, understands how it differs across customer populations and typologies, can identify possible automation bias or rubber-stamping, and can demonstrate that human oversight is still doing something real.

Agreement with the model is not proof of rubber-stamping. Disagreement with it is not proof of good judgement. This is the same discipline that applies to QA generally. Divergence is a signal to be explained through methodology and evidence, not a score to be celebrated or punished.

Models drift because the world moves

A model is approved against a set of assumptions.

Those assumptions do not freeze when the approval paper is signed. Customer behaviour changes. Transaction patterns change. Products, markets, sanctions measures, geopolitics, typologies, data quality and customer mix all change. Criminals adapt faster than validation cycles.

So before deployment the firm should decide which performance measures matter, what the baseline is, which cohorts and typologies are material, how re-testing will work, what the drift indicators are, and what triggers escalation, recalibration, restriction, suspension or reapproval where reapproval is required.

I would not prescribe quarterly validation, six-month samples, a five-point AUC trigger, a two-percentage-point deterioration threshold or five years of historical data. Those numbers get quoted as if they were requirements. They are not.

There is a harder problem underneath this. AML does not have clean ground truth. We rarely know with certainty which customers were laundering money and which were not. So drift cannot honestly be reduced to one performance number moving in one direction.

What firms can do is triangulate. Look at several indicators together and ask whether they are telling a consistent story.

No single one of those proves AML effectiveness. Taken together they usually tell you whether the assumptions behind approval still hold.

What the KRI is actually measuring

The KRI is not AUC, recall or an override rate by itself. The KRI is evidence that the model is moving away from the assumptions on which approval was granted.

The governance objective is modest and important. Identify meaningful deterioration, unexpected behaviour or changing assumptions early enough to investigate and respond.

Outcome testing beats vendor accuracy claims

A vendor tells me: “Our model reduces false positives by 60%.”

Fine.

My next question is: against what?

I care much more about what the model missed than how attractive the efficiency percentage looks on the sales slide.

Testing can draw on known historical cases, SAR and STR populations, typology sets, withheld validation populations, deliberately constructed challenge cases, higher-risk cohorts, complex ownership structures, relevant jurisdictions and emerging risk areas.

One qualification matters here. Historical SAR and STR decisions are not perfect ground truth. They contain human judgement made under the standards and pressures of the time. Standards move. Inconsistent historical decisions and bias can be learned and reproduced. Do not train the future against the weaknesses of the past and then present the result as validation.

The vendor does not inherit the firm's accountability

DORA is relevant here, and it is worth being accurate about it. It does not mean every AI failure is automatically a major ICT incident, and it does not require every AI problem to be reported to AMLA. What it does is put discipline around ICT dependency, resilience testing and third-party arrangements.

An AI dependency is still an ICT dependency.

A vendor SLA is not a substitute for a control fallback.

Explainability depends on the decision

Not every AI use case in a financial crime function is high-risk under the AI Act. Intended purpose and actual use decide that. GDPR implications likewise depend on the nature of the processing and how far the decision-making is automated. I would be careful with the claim that every AI-assisted AML decision creates a universal right to an explanation.

The word explainability is doing too much work in most discussions I sit in. It helps to separate three different things, because firms are often being criticised for lacking the first when what is actually missing is the second or the third.

I do not think an MLRO needs to rebuild a neural network on a whiteboard, or follow every interaction inside a graph model. Full technical interpretability may not be achievable at all for some architectures, and pretending otherwise wastes everyone's time.

That does not remove the obligation. The firm still has to be able to trace the decision and show who is accountable for it. Where technical interpretability is limited, I would expect the other two to be stronger, not weaker, and I would expect that limitation to be written down as an accepted limitation rather than left unmentioned.

Explainability is not the same thing as source-code disclosure. For an MLRO, it starts with being able to explain the control.

Put the risk trade-off beside the efficiency case

Suppose the business case says an AI tool will save €2 million in analyst cost. That may be a perfectly good reason to consider it.

The same paper should also say what human review disappears, which population is affected, what risk sits in that population, what the false-negative evidence shows, which cohorts retain mandatory human review, what compensating controls apply, what would trigger a view that performance has deteriorated, what residual risk is being accepted and who owns it.

An efficiency saving without an explicit risk transfer is not a complete business case. The Board should not be asked to approve €2 million of savings on one page and discover the control trade-off six months later in a risk paper. Put them beside each other.

Questions Boards should ask

  1. 01Which material financial-crime decisions does AI currently influence, and where does final human accountability sit?
  2. 02What evidence tells us that the AI-enabled control is performing within the limits we approved, and what would cause us to restrict or stop it?
  3. 03If the vendor and the model were unavailable tomorrow, could we reconstruct a material AI-influenced decision from our own records?
  4. 04What human review has been removed to create the efficiency benefit, and what residual risk have we accepted in return?

If those questions require the vendor to join the meeting before management can answer them, governance is already too far removed from the control.

Practical actions

How Claritas would assess an AI-enabled AML control

I would not start with the AI policy.

I would start with the work.

Show me where AI touches the customer journey. Show me what it can influence. Give me a deliberate sample of real cases. Show me the data it received. Show me what it produced. Show me what the analyst accepted, what they changed, what they rejected, and how the final conclusion was reached.

Then show me the model and version history, the performance MI, the drift indicators, the vendor dependencies, the stop conditions and whatever controls exist over unapproved use.

After that, I will read the policy.

The policy will tell me what should happen. The cases will tell me what actually did.

That is diagnostic and control-effectiveness work unless Claritas has explicitly been commissioned to provide formal Independent Assurance. Those are different pieces of work and I would not blur them.

What success looks like
  • 01The firm knows where AI is actually being used across financial-crime activity, including unapproved use that has been identified and brought under control.
  • 02Material AI use cases have a pre-deployment risk assessment, a clear purpose and clear decision rights.
  • 03Material AI-influenced decisions can be reconstructed from the firm's own case-level evidence without relying on the vendor to recreate the record.
  • 04Human oversight can be evidenced through reviewer behaviour rather than through a process diagram saying a human is involved.
  • 05Performance deterioration, population change and model drift lead to visible governance action where required.

Moving too slowly is also a risk

I want to be honest about the other side of this. Avoiding AI is not the risk-free option. Criminal actors are adopting the same technology, and they are not waiting for a governance committee. Synthetic identity material, fabricated source-of-wealth narratives and industrialised mule recruitment all get cheaper and more convincing when the other side automates.

So an AML function that treats every AI proposal as something to be delayed until it can be perfectly evidenced is accepting a risk too. It is simply accepting a quieter one.

The objective is not zero AI risk. It is understood AI risk.

Moving quickly without understanding what has been delegated to technology creates risk. Governance that makes responsible experimentation impossible creates a different one. Good governance should let a firm adopt sensibly, in controlled scope, with evidence, and stop if the evidence turns. That is not a brake on innovation. It is what makes innovation survivable.

Where I would land on it

I would approve AI in AML where it removes work that does not require human judgement. I would be much slower where it starts influencing the judgement itself.

And once it does, I want four things before I put my name behind it: the inputs, the output, the human intervention and the evidence behind the final decision.

If one of those disappears inside the model or inside the vendor, we have a governance problem.

I am not arguing for perfection here, and I would not want this read that way.

Residual risk that has been identified, sized, owned and accepted by the right people is governance. Residual risk nobody has articulated is not risk acceptance. It is an assumption waiting to be discovered by someone else.

That is my Empty Chair Test. Remove the technology from the room.

Can the firm still explain its own decision?

If not, I would want that gap understood and consciously accepted before the control goes anywhere near production. In most cases I have seen, once the gap is described plainly, nobody wants to accept it.

References
  • Regulation (EU) 2024/1624 (AMLR), including requirements relating to internal policies, controls and procedures and the assessment of risks associated with new or developing technology.
  • Regulation (EU) 2024/1689 (AI Act), classification by intended purpose and obligations attaching to that classification.
  • Regulation (EU) 2022/2554 (DORA), ICT risk management, resilience testing and third-party arrangements.
  • Regulation (EU) 2016/679 (GDPR), lawfulness of processing, international transfers and automated decision-making.
  • Official AMLA material on the future AML/CFT supervisory framework, including its role and the selection process for direct supervision.
  • European Banking Authority, banking sector risk assessment material published in June 2026, where relevant to technology adoption in the sector.
About the author
Everett Morgan
Founder & Principal Adviser, Claritas Risk Advisory

Everett has more than twenty years' experience in financial crime, AML governance, regulatory compliance and operational risk gained within Deutsche Bank, Morgan Stanley and BNP Paribas. He established Claritas Risk Advisory to provide smaller regulated financial institutions with experienced independent judgement, practical insight and proportionate recommendations.

Need an independent perspective?

Preparing for regulatory change starts with understanding where your organisation stands today.

If you would like to discuss your financial crime framework or explore how Claritas Risk Advisory can help, I would be pleased to arrange a confidential conversation.

Let's start with a conversation