Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation

📅 2025-04-21
📈 Citations: 0
Influential: 0
📄 PDF

career value

149K/year
🤖 AI Summary
Large language models (LLMs) exhibit opaque internal knowledge utilization, and their generated citations often lack faithfulness and credibility. Method: We propose the context-prior enhanced citation generation task, unifying the modeling of internal knowledge and externally retrieved knowledge in citation synthesis. We introduce RAEL—a novel retrieval-augmented, internal-knowledge-aware, and external-source-verified paradigm—and design INTRALIGN, a method integrating customized data synthesis with knowledge alignment to enable interpretable evaluation across utility, faithfulness, and credibility. Built upon a RAG framework, we establish a multidimensional credibility assessment system. Results: Our approach significantly outperforms mainstream baselines across diverse scenarios. Crucially, we quantitatively uncover the interactive effects of retrieval quality, question type, and model-inherent knowledge on citation credibility—establishing the first diagnostic, quantifiable citation evaluation paradigm for trustworthy AI.

Technology Category

Application Category

📝 Abstract
While hallucinations of large language models could been alleviated through retrieval-augmented generation and citation generation, how the model utilizes internal knowledge is still opaque, and the trustworthiness of its generated answers remains questionable. In this work, we introduce Context-Prior Augmented Citation Generation task, requiring models to generate citations considering both external and internal knowledge while providing trustworthy references, with 5 evaluation metrics focusing on 3 aspects: answer helpfulness, citation faithfulness, and trustworthiness. We introduce RAEL, the paradigm for our task, and also design INTRALIGN, an integrated method containing customary data generation and an alignment algorithm. Our experimental results show that our method achieves a better cross-scenario performance with regard to other baselines. Our extended experiments further reveal that retrieval quality, question types, and model knowledge have considerable influence on the trustworthiness in citation generation.
Problem

Research questions and friction points this paper is trying to address.

Enhancing transparency in LLMs' internal and external knowledge usage
Ensuring trustworthy citation generation in model responses
Evaluating answer helpfulness, citation faithfulness, and trustworthiness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-Prior Augmented Citation Generation task
RAEL paradigm for citation generation
INTRALIGN method with data alignment
🔎 Similar Papers
No similar papers found.
J
Jiajun Shen
University of Chinese Academy of Sciences
T
Tong Zhou
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences
Yubo Chen
Yubo Chen
Institute of Automation, Chinese Academy of Sciences
Natural Language ProcessingInformation ExtractionEvent ExtractionLarge Language Model
D
Delai Qiu
Unisound Al Technology Co,Ltd
S
Shengping Liu
Unisound Al Technology Co,Ltd
K
Kang Liu
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences
J
Jun Zhao
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences