When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence

πŸ“… 2026-07-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the tendency of large language models (LLMs) to overtrust derived measurements, mistakenly treating them as direct observations or applying them beyond their valid scope, thereby leading to erroneous judgments. The work introduces the first formal definition of the β€œDerived Feature Over-Trust” (DFOT) problem and proposes five quantifiable evaluation metrics. It further develops a general-purpose assessment and mitigation framework that does not rely on any specific reliability generator. By incorporating ECG as a privileged modality, the framework leverages ECG-to-PPG distillation, offline reference construction, and evidence-specific correction strategies. Evaluated on a locked test set of 187 patients, the approach significantly improves four key metrics by 1.82–6.69 percentage points, with confidence intervals excluding zero, demonstrating both efficacy and strong generalization.
πŸ“ Abstract
Derived measurements increasingly enter large language model (LLM) pipelines as direct facts despite their instance-dependent validity. We define derived-feature over-trust (DFOT) as the failure in which a downstream LLM assigns such a measurement the epistemic status of a direct fact or uses it outside its valid scope. Using physiological sensing as a case study, D1 tests acceptance of a PPG-derived rhythm contradicted by offline ECG, whereas D2 tests rejection of an offline-confirmed reliable PPG rhythm under misleading severe history. ECG supplies training supervision and offline reference construction but is never shown to the LLM. Five estimands quantify this chain: conflict over-trust rate (COTR) and context-induced error rate (CIR) characterize D1/D2; correct repair rate (CRR) measures frozen-error repair; evidence-specific repair margin (ESRM) contrasts matched and patient-disjoint shuffled evidence; and utility harm rate (UHR) measures unnecessary verification among HIGH-reliability cases used without verification at baseline. The framework does not depend on a particular reliability generator. We demonstrate it on 50,000 paired PPG-ECG records using ECG-to-PPG privileged distillation as an illustrative baseline and PPG-only inference. On a protocol-locked 187-patient test, the baseline improves four repair and specificity endpoints by 1.82-6.69 percentage points, with all paired confidence intervals excluding zero; UHR increases by 0.67 percentage points (95% CI: -0.4 to +1.7). DFOT provides a common evaluation target for stronger mitigation methods. The code is available at https://github.com/Zongheng-Guo/When-Derived-Measurements-Mislead.
Problem

Research questions and friction points this paper is trying to address.

derived measurements
large language models
over-trust
reliability
physiological sensing
Innovation

Methods, ideas, or system contributions that make the work stand out.

derived-feature over-trust
privileged-modality evidence
LLM reliability
conflict over-trust rate
evidence-specific repair
πŸ”Ž Similar Papers
2024-01-24Nature Machine IntelligenceCitations: 7