Local LLMs for AML: do they work in practice in 2026?
Three years on from our initial LLM AI benchmark for Financial Economic Crime, we return with a more operationally grounded test. Using 55,000+ sanctions entries across 12 sources, we evaluate whether local LLMs can reliably handle sanctions and transaction monitoring alert work in 2026. Tested models show meaningful variation in accuracy (51–81%), conclusion quality (71–93%), and speed. With the right tuning, the industry-standard 95% accuracy threshold appears within reach. Prerequisites include the use case is tightly scoped, the model well-selected, and the process thoughtfully implemented.






