🤖 AI Summary
This work addresses the lack of controllable evidence environments in existing fact verification benchmarks, which hinders robust evaluation of models under search engine optimization (SEO)-style evidence poisoning attacks. The authors propose the first multi-domain fact verification benchmark and evaluation framework that enables precise control over evidence sources and poisoning ratios, allowing systematic assessment of mainstream verification methods under adversarial retrieval conditions. Experimental results demonstrate that the proposed framework effectively uncovers robustness degradation and efficiency trade-offs that remain undetected when evaluating solely on clean data. These findings underscore the necessity of adversarial evaluation for fact verification systems and establish a new benchmark to guide the development of more robust verification approaches.
📝 Abstract
Large language models increasingly use search tools to retrieve up-to-date information, introducing a new attack surface in which retrieved documents can be manipulated. This risk is amplified by the development of generative engine optimization, which can make selected content more likely to be retrieved, cited, and adopted by models. Existing fact-verification benchmarks and evaluation frameworks do not provide the controlled evidence environments needed to assess robustness against GEO poisoning. We therefore propose GPE, which consists of a multi-domain fact-verification benchmark and an evaluation framework for controlling evidence sources and poisoning ratios. Experiments across multiple verification methods and poisoning attacks demonstrate that GPE exposes robustness degradation and efficiency trade-offs that cannot be observed through clean evaluation alone, confirming the need to evaluate fact verification under adversarial evidence environments.