TokenShapley: Token Level Context Attribution with Shapley Value

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing context learning attribution methods operate only at the sentence level, failing to support fine-grained, token-level provenance—particularly for critical lexical units such as numbers, years, and proper nouns. To address this limitation, we propose TokenShapley, the first method to apply Shapley values for token-level attribution in large language model (LLM) response generation. TokenShapley integrates a KNN-enhanced retrieval mechanism to enable efficient context matching within a pre-constructed data repository. By doing so, it overcomes the granularity bottleneck of conventional attribution approaches and delivers interpretable, high-precision quantification of individual token contributions. Evaluated across four benchmark tasks, TokenShapley significantly outperforms state-of-the-art baselines, achieving 11–23% higher token-level attribution accuracy. This work establishes a novel paradigm for trustworthy LLM reasoning and fine-grained溯源 analysis.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Planning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Large language models (LLMs) demonstrate strong capabilities in in-context learning, but verifying the correctness of their generated responses remains a challenge. Prior work has explored attribution at the sentence level, but these methods fall short when users seek attribution for specific keywords within the response, such as numbers, years, or names. To address this limitation, we propose TokenShapley, a novel token-level attribution method that combines Shapley value-based data attribution with KNN-based retrieval techniques inspired by recent advances in KNN-augmented LLMs. By leveraging a precomputed datastore for contextual retrieval and computing Shapley values to quantify token importance, TokenShapley provides a fine-grained data attribution approach. Extensive evaluations on four benchmarks show that TokenShapley outperforms state-of-the-art baselines in token-level attribution, achieving an 11-23% improvement in accuracy.
Problem

Research questions and friction points this paper is trying to address.

Attributing specific keywords in LLM responses
Improving token-level context attribution accuracy
Combining Shapley values with KNN retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-level Shapley value attribution
KNN-based retrieval techniques
Precomputed datastore for contextual retrieval
🔎 Similar Papers
No similar papers found.