🤖 AI Summary
This study addresses the semantic deficiency in financial fraud data caused by privacy restrictions and the insufficient behavioral authenticity of synthetic data. We propose a transaction behavior-driven multi-agent semantic enhancement framework that introduces a novel behavior-anchored multi-agent consistency refinement mechanism, integrating large language model reasoning with statistical fidelity verification. This approach generates MS-FFSD, a high-fidelity multimodal financial fraud dataset that preserves authentic transaction behaviors while enriching structured and textual semantics. Experimental results demonstrate that such semantic enrichment significantly improves both fraud detection modeling accuracy and the contextual reasoning capabilities of large models. Furthermore, the framework exhibits strong generalizability, effectively bridging critical data gaps in this domain.
📝 Abstract
In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However, public real-world financial datasets often lack rich semantics due to privacy constraints. Consequently, synthetic datasets incorporate generated semantics, but at the cost of behavioral realism; textual descriptions for contextual reasoning remain scarce. We address this gap through a semantic enrichment framework grounded in original transaction behavior to simulate multimodal financial data. We (1) propose a multi-agent semantic enrichment framework that generates interpretable financial semantics grounded in transaction behavior through role-specialized agents and consistency refinement, and (2) newly contribute a valuable multimodal financial fraud dataset, MS-FFSD, enriched with structured semantics and textual semantics while preserving real-data-grounded transaction behavior. Furthermore, we systematically analyze the quality and utility of semantic enrichment. Results demonstrate statistical fidelity and framework generalizability, while showing that richer semantics benefit fraud modeling and context-aware LLM reasoning. Overall, this work advances multimodal financial fraud research and bridges emerging LLM and multi-agent capabilities with operational anti-fraud practice. The framework and dataset are released at https://github.com/AI4Risk/MS-FFSD.