Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在以检索为主的抽取式问答中,使用预设负例的方法来解决模型自信度信号不足的问题,但实验表明该方法未能显著提升性能。
📝 Abstract
In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy -- confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields a model whose own confidence is a tempting control signal: it could decide which queries warrant further adaptation, and which answers to trust. We evaluate both uses under criteria fixed before the runs were executed, across four 7-9B model families whose adaptation moved closed-book F1 by at most +0.03, and both fail: a distillation trigger on all four families, under its pre-specified three-step transfer budget, and a routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy under every correctness criterion we test, leaving routers no meaningful gain. The sequence-likelihood signal is insufficient relative to that mode -- area under the receiver operating characteristic curve 0.65-0.81 under the registered criterion -- before adaptation as well as after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. And the finer diagnostics depend on the correctness criterion and on answer length; on the three adapted combinations where we could test it, selector ablations show no statistically detectable downstream benefit from the confidence term on any seed; on Gemma, removing it changes the selector from failing to passing both registered criteria. The usable product is a set of pre-specified negatives with their dependencies made explicit.
Problem

Research questions and friction points this paper is trying to address.

extractive QA
retrieval
confidence-driven mechanisms
fine-tuning
language model
Innovation

Methods, ideas, or system contributions that make the work stand out.

intrinsic sequence-likelihood confidence
pre-specified negatives
retrieval-dominated QA
confidence-driven mechanisms
fine-tuning
Gunwoo Lee
Gunwoo Lee
Large-scale AI Research Center, Korea Institute of Science and Technology Information (KISTI), 245 Daehak-ro, Yuseong-gu, Daejeon, 34141, Republic of Korea.
C
Changmin Sung
Large-scale AI Research Center, Korea Institute of Science and Technology Information (KISTI), 245 Daehak-ro, Yuseong-gu, Daejeon, 34141, Republic of Korea.
S
Sang-Hwan Gwak
Large-scale AI Research Center, Korea Institute of Science and Technology Information (KISTI), 245 Daehak-ro, Yuseong-gu, Daejeon, 34141, Republic of Korea.
I
Ina Kim
Large-scale AI Research Center, Korea Institute of Science and Technology Information (KISTI), 245 Daehak-ro, Yuseong-gu, Daejeon, 34141, Republic of Korea.
J
Ji-Young Choi
Large-scale AI Research Center, Korea Institute of Science and Technology Information (KISTI), 245 Daehak-ro, Yuseong-gu, Daejeon, 34141, Republic of Korea.
K
Kyong-Ha Lee
Large-scale AI Research Center, Korea Institute of Science and Technology Information (KISTI), 245 Daehak-ro, Yuseong-gu, Daejeon, 34141, Republic of Korea.; Department of Applied AI, University of Science and Technology (UST), 217 Gajeong-ro, Yuseong-gu, Daejeon, 34113, Republic of Korea.