🤖 AI Summary
Large language models (LLMs) frequently generate “hallucinations”—plausible yet factually incorrect statements—posing significant risks in high-stakes domains such as clinical decision-making. Existing evaluation metrics struggle to jointly ensure factual consistency and interpretability, hindering precise error diagnosis and correction. To address this, we propose FACT-DECOMP: a pattern-agnostic, semantic decomposition–based evaluation framework that quantifies factual consistency at fine-grained, interpretable levels. It comprises atomic fact extraction, semantic alignment modeling, weighted consistency scoring, and dynamic complexity control. Evaluated on both general-purpose and clinical benchmarks, FACT-DECOMP consistently outperforms state-of-the-art metrics (e.g., FactScore, FEVERScore) in accuracy and robustness, while supporting unified assessment across open-domain and domain-specific texts. The implementation is publicly available, providing a reproducible, debuggable foundation for developing fact-aware LLMs.
📝 Abstract
Large Language Models have significantly advanced natural language processing tasks, but remain prone to generating incorrect or misleading but plausible arguments. This issue, known as hallucination, is particularly concerning in high-stakes domains like clinical applications, where factual inaccuracies can have severe consequences. Existing evaluation metrics fail to adequately assess factual consistency and lack interpretability, making diagnosing and mitigating errors difficult. We propose an interpretable framework for factual consistency assessment for in-domain and open-domain texts to address these limitations. Our approach decomposes text into atomic facts and introduces a flexible, schema-free methodology. Unlike previous methods with an absolute metric, we incorporate a weighted metric to enhance factual evaluation. Additionally, we propose a mechanism to control assessment complexity in intricate domains. We benchmark our approach on popular general and clinical datasets and release our code to support fact-aware model training in future research.