arXiv Science⌕ Search

arXiv · 2609.37155

Explainable and Trustworthy AI for Anti-Money Laundering: A Graph-based Hybrid Framework for Real-World Financial Crime Detection

Abstract

Purpose: Money laundering threatens financial systems, while rule-based monitoring suffers from high false-positive rates and limited ability to capture relational transaction patterns. This study proposes an explainable graph-based framework for anti-money laundering (AML) detection that jointly addresses predictive performance, explanation faithfulness, and uncertainty calibration. Methods: Using the IBM Transactions for Anti-Money Laundering (HI-Small) benchmark, comprising 5,078,345 transactions among 518,573 accounts with 0.10% illicit transactions, a directed attributed graph with 4,487,133 edges was constructed using temporal, leakage-safe partitioning. Three GATv2 architectures, BASE, BASE-Large, and IMPROVED, incorporating bidirectional message passing, edge updates, and port-aware features, were compared with XGBoost and Random Forest using identical features. The best model was evaluated using GNNExplainer against documented typologies and Mondrian conformal prediction for uncertainty calibration. Results: IMPROVED achieved the strongest performance (AUPRC = 64.85%, Best F1 = 68.42%), exceeding BASE-Large by 30.65 AUPRC points and XGBoost (AUPRC = 38.46%) by 26.4 points. The proposed mechanisms contributed more than capacity scaling alone. GNNExplainer recovered documented typologies with higher fidelity than attention-weight and random baselines (mean Jaccard overlap: 21.4% vs. 1.0%, p < 0.001). Mondrian conformal prediction achieved coverage close to the 90% nominal target with an average prediction-set size near one. Conclusion: Explicit transaction topology modeling substantially improves AML detection, while faithful explanations and calibrated uncertainty support interpretable, human-in-the-loop compliance decision-making.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ali Shahbazi, Faraz Sasani, Arshia Hossein zadeh, Sheyda safaeimoradi, Hossein Najafzadeh. 2026-09-29. Explainable and Trustworthy AI for Anti-Money Laundering: A Graph-based Hybrid Framework for Real-World Financial Crime Detection. https://arxiv.org/abs/2609.37155

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Trie Constraints and Hierarchy-Aware Semantic Alignment for HS Code Prediction with Small Language Models

Harmonized System (HS) code prediction (HSP) from commodity text is essential to international trade, and its importance continues to grow in port logistics. Recently, large language models (LLMs) have been actively investigated for this task, owing especially to their strong language-understanding capabilities. However, their high computational cost limits deployment in constrained environments such as container terminals. Small language models (SLMs) offer a practical alternative, but their smaller scale makes them prone to generating invalid HS codes and to overlooking the hierarchical semantics between commodity text and HS codes. To address these limitations, this study proposes TRIE-HSA, which combines trie-constrained token prediction with hierarchy-aware semantic alignment (HSA). This framework constrains the SLM to predict only valid digits under the HS taxonomy and aligns commodity text representations with the hierarchical structure of HS codes. In extensive experiments on data collected from an operational container terminal, TRIE-HSA improved average HS6 accuracy by 49.96% over zero-shot inference and exceeded the strongest task-specific benchmark by 11.94%. These results demonstrate that accurate and structurally valid HSP is achievable with fewer than 10 billion parameters. Therefore, TRIE-HSA offers a practical basis for deployment of HSP in port logistics operations that cannot support large scale LLMs.

cs.CE↗

Adaptive Parallel-in-Time Integration with Dynamic Resource Management

As computational resources continue to grow, the strong-scaling limitations of spatial parallelism motivate the pursuit of additional concurrency in the temporal dimension, particularly for applications with hard time constraints, such as weather and climate simulations. The Parallel Full Approximation Scheme in Space and Time (PFASST) is a parallel-in-time method based on Spectral Deferred Corrections (SDC). It computes multiple timesteps concurrently by coupling fine- and coarse-grid SDC sweeps using multigrid Full Approximation Scheme (FAS) corrections. However, PFASST's convergence is often problem-dependent, demanding a variable number of parallel timesteps and, hence, computing resources at different times throughout the simulation. Dynamic Resource Management (DRM) provides a remedy for this challenge by enabling the adaptive adjustment of computational resources and algorithmic parameters at runtime. In this work, we present our novel approach to extending PFASST with DRM, which enables (a) dynamic adaptation of computing resources, (b) adaptive selection of the number of PFASST iterations based on local convergence behavior, and (c) coupling of these two adaptations into a single resizing strategy. With this approach, we demonstrate for the first time that optimal configurations can be identified in real time for each application, rather than relying on static allocation. Furthermore, we show that convergence-informed tuning of PFASST improves resource utilization and convergence efficiency.

cs.CE↗

A Three-Layer Framework for Measuring Names and Its Census Application on a Token Launchpad

Asset names influence market behavior, yet standardized name measurement remains lacking. Existing processing fluency measures focus mainly on alphabetic languages and are unsuitable for Chinese names. Cultural meanings usually require manual coding, limiting large-scale analysis, while name competition through reuse and semantic crowding remains underexplored. This study constructs a dataset of 513,647 naming attempts from the Four token launchpad on BNB Chain between February and June 2026 and proposes a three-layer framework for name measurement. The form layer measures linguistic fluency using 38 Chinese-oriented features. The reference layer captures cultural meanings through human coding and large language model expansion with reliability evaluation. The relation layer measures name reuse, semantic crowding, and lexical variation. The three layers are largely independent, with correlations below 0.11. Census analysis reveals that name diversity follows Heaps' law, new-name adoption declines over time, name reuse shows heavy-tailed patterns, and repeated naming occurs at distinct creator-level and cross-creator time scales. Cultural events also trigger rapid naming responses. The framework, annotated dataset, and code are released to support scalable analysis of naming behavior in digital markets and other naming environments.

cs.CE↗