evidence-aligned qa filtering

Designs and implements automated filters that evaluate question–answer pairs for alignment with supporting evidence, using QA-based heuristics and support-scoring to flag or remove answers likely to be rejected; includes mechanisms to align filter tests with verifier evidence and to selectively skip verifier checks when evidence is sufficient or inapplicable.

evidence-alignedqafiltering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

When large language models (LLMs) serve as automated judges for question-answering evaluation, their scoring can be distorted by internal belief conflicts when the reference answer contradicts their pretrained knowledge. This work proposes a “reference-swapping” framework that constructs aligned pairs of candidate answers via entity substitution, enabling systematic investigation of LLM judges’ behavior under controlled reference–belief conflicts. Experiments demonstrate that the scoring reliability of mainstream LLM judges significantly degrades in such settings, and existing prompting strategies fail to adequately mitigate this issue. This study is the first to identify and quantify this failure mode, revealing a fundamental limitation in current automatic evaluation methodologies.

evaluation fidelityknowledge conflictLLM-as-a-judge

This study addresses the instability and poor interpretability of short-answer visual question answering (VQA) benchmark evaluations, which often misclassify semantically correct answers as errors due to overreliance on superficial string matching. Leveraging a high-precision (97.6%) human-validated semantic judgment protocol, the authors conduct a systematic audit of over 37k official errors from six multimodal models across six benchmarks, revealing that nearly half of these “errors” are in fact semantically accurate but differ in surface form. Through text-only model replication, deterministic CPU-based repair contracts, and answer-type diagnostics, the work demonstrates that evaluation bias stems from the scorer’s preference for lexical form over meaning, and shows that simple prompting or contextual fine-tuning can substantially improve scoring stability. The study advocates semantic auditing and answer-type analysis as essential complements to standard VQA benchmark evaluation.

evaluator-dependent instabilitymultimodal benchmarkssemantic correctness

This work addresses the challenge that large language models often produce responses in high-stakes scenarios whose correctness is difficult for users to verify based on the provided reasoning. The paper introduces “error verifiability” as a novel quality dimension distinct from accuracy and proposes a balanced metric, \( v_{\text{bal}} \), to quantify how effectively a model’s rationale aids human judgment. To enhance verifiability, the authors develop two external-information-augmented rephrasing methods: Reflect-and-Rephrase (RR) for mathematical reasoning and Oracle-Rephrase (OR) for factual question answering. Experimental results demonstrate that conventional training and scaling approaches fail to improve verifiability, whereas RR and OR substantially increase human accuracy in judging answer correctness.

error verifiabilityhuman evaluationjustifications

Hot Scholars

TG

Tanya Goyal

Cornell University
Natural Language Processing
SA

Sophia Ananiadou

Professor, Computer Science, Manchester University, National Centre for Text Mining
Natural Language ProcessingText MiningComputational LinguisticsArtificial Intelligence
CJ

Christopher J. MacLellan

Assistant Professor, Georgia Institute of Technology
Cognitive SystemsArtificial Intelligence in EducationHuman-AI TeamingConcept Formation
MS

Mubarak Shah

Trustee Chair Professor of Computer Science, University of Central Florida
Computer Vision
KM

Karthik Mohan

University of Toronto
Machine LearningArtificial IntelligenceComputer Science