
Research Article
Temporal-Structural Stress Testing and Cross-Granularity Robustness Evaluation of Graph AI for Cryptocurrency AML Detection
@ARTICLE{10.4108/airo.13658, author={Sadia Akter and Sabab Islam and Md Faysal Ahmed and Md Hossain Jamil and Md Fokrul Islam Khan and Partha Singha and Abu Kowshir Bitto and Sanim Yousuf Fahim }, title={Temporal-Structural Stress Testing and Cross-Granularity Robustness Evaluation of Graph AI for Cryptocurrency AML Detection}, journal={EAI Endorsed Transactions on AI and Robotics}, volume={5}, number={1}, publisher={EAI}, journal_a={AIRO}, year={2026}, month={9}, keywords={cryptocurrency AML, graph AI, graph neural networks, financial crime, temporal validation}, doi={10.4108/airo.13658} }- Sadia Akter
Sabab Islam
Md Faysal Ahmed
Md Hossain Jamil
Md Fokrul Islam Khan
Partha Singha
Abu Kowshir Bitto
Sanim Yousuf Fahim
Year: 2026
Temporal-Structural Stress Testing and Cross-Granularity Robustness Evaluation of Graph AI for Cryptocurrency AML Detection
AIRO
EAI
DOI: 10.4108/airo.13658
Abstract
Cryptocurrency anti-money-laundering (AML) is typically framed as a graph-learning problem, since suspicious value flows rarely appear as isolated records. This study examines whether graph AI architectures retain an advantage over strong non-GNN baselines when evaluation is temporal, structurally explicit, and aligned with investigator triage. We answer that question using two public Bitcoin AML datasets: Elliptic, a transaction- node dataset, and Elliptic2, a subgraph-level dataset. Our comparison spans classical models, tree ensembles, graph-derived features, graph embeddings, neural baselines, and GNN variants, evaluated across AUPRC, AUROC, accuracy, precision and recall at an investigation budget, calibration, operating points, and bootstrap uncertainty. On the temporal Elliptic test period, Extra Trees achieves AUPRC 0.570 and AUROC 0.869, while GraphSAGE, the strongest standard GNN baseline, reaches AUPRC 0.311. A strict-temporal GraphSAGE variant improves to AUPRC 0.381 but still falls behind the leading tree ensemble model. The paired bootstrap difference between Extra Trees and GraphSAGE is 0.259 AUPRC, with a 95% interval of [0.222, 0.297]. On the full Elliptic2 subgraph benchmark, Random Forest achieves AUPRC 0.494 and AUROC 0.923 in the main split and remains the strongest model by mean AUPRC across repeated splits. These findings indicate that graph AI for cryptocurrency AML should be stress-tested against capable tree ensemble models and operational metrics before architectural complexity is treated as deployment evidence.
Copyright © 2026 Sadia Akter et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.


