AIC CTU@FEVER 8: On-premise fact checking through long context RAG

📅 2025-08-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying high-accuracy fact-checking on resource-constrained edge devices—specifically, a single NVIDIA A10 GPU (23 GB VRAM) with a strict 60-second inference latency budget. Method: We propose a lightweight, end-to-end, locally deployable two-stage framework that innovatively integrates long-context retrieval-augmented generation (RAG) with a fine-grained natural language inference (NLI) module. By optimizing the retrieval-verification synergy, our approach significantly improves evidence coverage and logical reasoning capability within limited context windows. Contribution/Results: Our system achieves first place in the FEVER 8 shared task and sets a new state-of-the-art (SOTA) on the Ev2R benchmark. It meets both hardware and latency constraints in practice, demonstrating the feasibility of accurate fact-checking on edge-grade hardware. Key contributions include: (1) a resource-aware RAG architecture design; (2) a low-overhead strategy for long-context modeling; and (3) an open-source, reproducible lightweight fact-checking system.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Knowledge Representation and Reasoning: Reasoning with BeliefsMachine Learning: Learning on the Edge & Model Compression

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsResponsible Web: Human-perceived consequences of algorithmic deployment on the webGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
In this paper, we present our fact-checking pipeline which has scored first in FEVER 8 shared task. Our fact-checking system is a simple two-step RAG pipeline based on our last year's submission. We show how the pipeline can be redeployed on-premise, achieving state-of-the-art fact-checking performance (in sense of Ev2R test-score), even under the constraint of a single NVidia A10 GPU, 23GB of graphical memory and 60s running time per claim.
Problem

Research questions and friction points this paper is trying to address.

Develop on-premise fact-checking pipeline
Achieve state-of-the-art performance with constraints
Optimize RAG pipeline for limited GPU resources
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-premise fact checking pipeline
Two-step RAG for efficiency
Optimized for limited GPU resources
🔎 Similar Papers
No similar papers found.