Score
Designs and implements mechanisms that dynamically gate or weight corrective signals, secondary evidence, or model updates according to assessed trust or epistemic uncertainty at per-sample, per-variable, per-channel, or hierarchical (multi-level) scopes. Builds the trust-estimation metrics and integrates them into control or fusion pathways so corrections are attenuated when cues are unreliable and amplified when confidence is high, and analyzes how those gating decisions affect downstream performance.
研究测试了大语言模型作为裁判时信任评分与真实性判断之间的分离性,通过改变信息来源进行压力测试,发现两者并不完全独立。
本文提出CRiDiT测试平台,通过Dempster-Shafer理论和主观逻辑等方法解决AI系统中人机信任校准问题。
Existing research lacks standardized, context-aware human-AI trust metrics, hindering clear distinction between trust formation and decision execution. Method: We propose the first dynamic trust calibration framework, modeling human adaptive trust adjustment toward AI outputs via contextual bandits to explicitly decouple opinion formation from final decision-making, and introducing a generalizable quantitative metric system. Contribution/Results: Evaluated across three cross-domain datasets—including high-stakes medical and judicial domains—the framework significantly improves human-AI collaborative decision-making performance, yielding 10–38% reward gains over baselines. It establishes a novel, interpretable, and reproducible paradigm for trust calibration in safety-critical applications, enabling principled, context-sensitive human-AI coordination.
This work addresses the vulnerability of large language models acting as autonomous agents in complex tool-augmented environments, where contextual contamination and over-reasoning often arise due to the absence of effective mechanisms for evaluating both their own capabilities and the reliability of external tools. To tackle this, the paper introduces the MESA-S framework, which, for the first time in a single-agent setting, decouples skill utility awareness from execution. By integrating metacognitive skill cards, delayed process probes, and a dual-dimensional confidence model, MESA-S computationally instantiates human-like cognitive control mechanisms—including delayed evaluation, cognitive vigilance, and proximal offloading—and formalizes trust provenance. Experimental results demonstrate that this approach effectively mitigates supply-chain vulnerabilities, curbs confidence inflation, reduces redundant reasoning, and significantly enhances agent behavioral reliability.
This work addresses the problem of trust calibration in autonomous agents—specifically, how an agent should dynamically decide whether to act independently or seek human approval when using automated tools. The paper formalizes this challenge as a preference learning task for the first time. It introduces a policy gateway that maintains a Gaussian process posterior over the human’s risk tolerance function, employing a probit likelihood and an approximate Gaussian process classification model to infer preferences from binary approve/reject feedback. The agent actively queries human input at points of highest uncertainty, thereby establishing a three-region decision mechanism: “allow,” “block,” and “ask.” This approach extends the applicability of preference-based Bayesian optimization and achieves sample-efficient trust calibration, accurately partitioning the action space while substantially reducing unnecessary human interventions.
This study addresses the security vulnerability arising from erroneous trust in large language model (LLM) assistants when they fail to verify user intent. To mitigate this, the authors construct contrastive dialogue datasets and identify linear directions within the activation space, enabling causal intervention on the model’s trust decisions through activation steering matrices while keeping parameters frozen. This work provides the first demonstration that LLM trust behavior can be controlled both monotonically and causally via such a mechanism. Furthermore, it achieves bidirectional trust regulation across multiple model families, effectively mitigating critical security threats, including harmful requests and prompt injection attacks.
This study addresses the paradox in multi-agent LLM systems wherein high mutual trust enhances performance yet exacerbates vulnerability. We present the first operationalization of a hierarchical trust model by constructing a five-layer trust stack. Through generalized-mean composite trust aggregation, regret-minimization-based online weight adaptation, and short-term revocable capability tokens, the framework achieves cross-layer coordination and dynamically trustworthy capability delegation. Theoretically, we prove that this approach breaks the trust–vulnerability paradox. Empirically, experiments confirm the validity of composite trust bounds, yielding a fixed-point error below 0.007, and demonstrate that compromised premise layers trigger automatic revocation within a few interactions.
This study addresses the "shortcut learning" bias in fact-checking agents, which often over-rely on source credibility labels while neglecting evidence content. To quantify this bias, we introduce TrustSwap, a novel multi-channel counterfactual testing framework, and propose a Trust-Swap Augmentation (TSA) strategy. TSA mitigates label dependency by integrating label-swapping-based GRPO data augmentation with reinforcement learning to fine-tune retrieval-augmented generation models. Experiments demonstrate that our approach significantly reduces verdict flip rates in 4B-parameter models while preserving accuracy and confidence calibration. Outperforming conventional reward schemes, this work establishes a new paradigm for enhancing the robustness of fact-checking systems.
本文通过整合可信AI治理、代理安全等方法,提出信任即服务(TaaS)来解决跨组织边界行动的信任评估问题。
研究通过设计非指导性苏格拉底对话的会话代理CASELy来校准用户信任,避免用户对AI系统的过度或不足依赖。