Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
Current evaluations of financial large language models (LLMs) rely excessively on static benchmarks and lack comprehensive validation across the full system stack. This work proposes the first LLM full-stack verification framework tailored to financial scenarios, encompassing data, model, retrieval, generation, agent behavior, governance, and deployment layers, advocating for verification as an ongoing engineering practice. By introducing a multi-rater LLM-as-a-judge mechanism integrated with scoring rubrics, consistency checks, and auditability, the framework uncovers system failure modes that static benchmarks fail to capture. The study defines critical failure types and advances novel directions—including system-aware benchmarks, agent trajectory validation, rater alignment protocols, and lifecycle-oriented verification standards—thereby shifting the evaluation paradigm from score-driven metrics toward evidence-based readiness for real-world decision-making.