Towards Robust Numerical Claim Verification

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high vulnerability of large language models to minute numerical perturbations in numerical claim verification. Leveraging the Qwen3 model series, we employ parameter-efficient fine-tuning (PEFT) to conduct adversarial training on numerically perturbed samples. Our experiments demonstrate that the resulting fine-tuned small-scale models surpass frontier systems such as GPT-5.4 Pro, achieving 98.7% accuracy under label-flipping perturbations. Furthermore, these models exhibit strong generalization to unseen perturbation types and deliver robust cross-lingual and cross-domain transfer performance even without access to target-domain data.
πŸ“ Abstract
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Using parameter-efficient fine-tuning, small Qwen3 models (0.6B$\unicode{x2013}$8B) reach 98.7% accuracy on label-flipping perturbations, outperforming larger zero-shot models and frontier systems (GPT-5.4 Pro (74.0%) and Gemini 2.5 Flash (73.9%)). The gains generalise to unseen perturbation types, indicating robust numerical decision boundaries rather than memorised edits. Robustness also transfers without target-domain data, significantly improving cross-lingual performance in Spanish. We further show that the same fine-tuning recipe confers robustness to evidence-side perturbations, using the VitaminC dataset.
Problem

Research questions and friction points this paper is trying to address.

Claim Verification
Numerical Reasoning
Robustness
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Numerical Claim Verification
Adversarial Fine-tuning
Parameter-Efficient Fine-Tuning
Robustness Generalization
Cross-lingual Transfer
πŸ”Ž Similar Papers