I recently joined the ACAMS New Jersey chapter for a discussion on transitioning from threshold-based transaction monitoring (TM) to more advanced, AI-driven systems. Alongside Samrat Jain of PwC, Murat Dagli of Remitly, and our moderator Chris Phillips we dove into an ongoing debate among financial crime professionals: what is the potential for AI in TM versus what are teams currently achieving with deployments.
Not all AI is equal in AML
Most institutions using AI in transaction monitoring today are using it for efficiency: helping investigators get through their alert queue faster, gathering context on a customer, or drafting the first version of a Suspicious Activity Report (SAR) narrative for a human to finalize. That work is genuinely valuable. Anyone who has sat in the investigator's seat knows how much of the day goes to assembling information before you can even start making decisions. But in every one of those use cases, the detection has already happened. The rules fired, the alert exists, and the AI is just helping analysts to process it. It isn't finding anything the rules didn't flag.
A smaller but growing group is putting AI on the detection side itself, particularly since the U.S. Department of the Treasury’s Financial Crimes Enforcement Network (FinCEN) issued a Notice of Proposed Rulemaking (NPRM) in April, intended to fundamentally reform financial anti-money laundering and counter-terrorist financing (AML/CFT) programs under the Bank Secrecy Act (BSA). What this looks like in practice is models running alongside the rules, learning what normal looks like for each customer, and surfacing activity the rules were never designed to catch. That's a fundamentally different use of the technology, and it is where regulators are investing effort to drive improved outcomes.
It's also worth pulling apart what we mean by AI in the first place: traditional machine learning that learns a customer's normal patterns and scores alerts based on which ones have historically been real; generative AI doing the drafting and summarizing; and, increasingly, agentic AI, where a system is given a goal and decides its own steps to get there without a human holding its hand. The governance and risk each of these needs is completely different, so when someone tells you their program “uses AI,” the first question should be: which one?
FinCEN NPRM enforcement focus
FinCEN named AI directly as a factor it will weigh when deciding whether to bring an enforcement action. For years, our industry has asked how much time AI can save us, or whether it lets us repurpose headcount. The NPRM asks a different question: how much crime does it actually help you identify? Most AI running in TM programs today was built to answer the first question. After more than a decade in this space, I think the shift toward the second is a welcome one — even if there are still plenty of open questions about how it gets audited and enforced in practice. This shift from FinCEN echoes the focus by leading global regulators on outcomes as opposed to check-the-box compliance.
Shifting SARs from defense to precision
I've written probably far too many “super SARs” in my day — SARs so overloaded with every recipient and every transaction that they bury the actual story. There's a real inclination in this industry to file defensively: when you don't have a clear explanation, the instinct is to put everything in the SAR and let law enforcement figure out what is pertinent. In reality, this overloads law enforcement and lets criminal activity hide in the noise. Great SARs are precise, they tell the story the investigation uncovered, with a clear narrative and the evidence that supports it, rather than the whole kitchen sink. Narrative explanations from AI audit trails can support improved SAR construction and precision to support improved crime prevention by law enforcement, alongside enhanced compliance for financial crime teams.
Build vs buy: calculating the total lifecycle cost of ownership
Whether building in-house makes sense really depends on existing resources and institutional knowledge: real data scientists and machine learning engineers, and a willingness to treat these tools as a product you run and maintain forever. Some banks have done it well. But there's a reason banks don't build their own core banking systems or case management tools from scratch. A detection system isn't a project you finish — it's something you run for as long as you're monitoring transactions, through every new product, every data change, every sanctions list update. Once a CTO prices all of that in, building is often the more expensive road.
What gets underweighted in most build business cases is what you give up on the detection side by going it alone. A vendor sees typologies play out across hundreds of institutions, jurisdictions and payment rails; an in-house team sees its own book. When a new laundering pattern surfaces in one market, that learning gets folded back into the models everyone runs — the consortium effect. It is difficult to replicate from a single institution's data, however good your data scientists are.
It also isn't a binary choice. The institutions that have handled this well treated it as a set of questions rather than one decision: how mature is our technology stack, how heavy is our regulatory burden, do we have the machine learning and model operations skills to run this in production for the next decade, and how quickly do we need to scale? Answer those honestly and the pattern is fairly consistent: buy where speed, cost-efficiency, and regulatory readiness are what you are optimizing for; and build where you have genuine internal capability and a strategic reason to differentiate. In practice, most programs land somewhere in between: buy the engine, and keep the risk logic and tuning in-house, where your institutional knowledge resides.
And when you do price it, price the whole lifecycle rather than the build. The costs that catch teams out arrive after go-live: model validation and revalidation, monitoring for drift, retraining as products and customer bases change, and the documentation you need every time an examiner asks why a model reached a particular decision. None of that is optional, and none of it stops.
For a detailed build vs buy analysis, download the whitepaper: How to calculate the total cost of ownership for AML transaction monitoring
Making improved outcomes the goal
If the only thing AI is doing in your AML program is churning through the same false positives slightly faster than before, then it hasn't actually improved detection or reduced risk exposure — it’s simply processing the same alert queue more quickly. It's tempting to point to close rates and call that progress, but if your false positive rate is still north of 90%, there's likely real financial crime being left on the table and unknown false negatives that represent real risk exposure for your organization.
It's also worth setting expectations early: when models start running alongside rules, alert volumes typically go up in year one, not down, because the model is surfacing categories of activity your investigators haven't handled before. And workload will likely increase, meaning it is unwise to promise the board cost reduction in three months. Regulators have said explicitly — across FinCEN, Treasury, the SEC, and others — that risk-based reallocation isn't meant to translate into smaller teams, and that improving detection shouldn't be treated as an admission that a program was deficient before.
Humans in the AML decision loop
On whether AI should be allowed to auto-close alerts or file SARs end-to-end, my answer is still no, for the foreseeable future. Judgment can't be outsourced: someone has to be accountable for the decision, and I don't think most people are ready to put a model's name on that or regulators ready to permit it. But there's a follow-on question worth sitting with: if an alert is low-risk enough that we'd be comfortable letting a model close it automatically, why is it being generated in the first place? Usually, the better fix isn't more automation at the closing end — it's tuning detection so the alerts your system produces are the ones worth a real investigator's time.
The role of rules in AML TM
Rules aren't going anywhere, and I don't think they should. The more useful frame is a feedback loop: models surface what rules can't catch, and once a pattern proves itself out through real, meaningful SARs, it gets built back into an explainable rule. Over time, your rule library becomes a reflection of everything the models have identified. That's the version of AI in transaction monitoring I think holds up to regulatory scrutiny — and the one worth building toward.











