Dual-Frontier: When Can an Agent Trust Its World Model?

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了代理信任其世界模型的问题,通过Dual-Frontier方法确保决策优势超过误差界限时才采用世界模型指导的决策。
📝 Abstract
Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not reveal whether the agent's decision rule or the world model caused the loss. We formalize this failure-attribution problem as a counterfactual decomposition of return loss and prove that its components are not identifiable from passive interaction, even for finite-horizon planners. This obstruction motivates Dual-Frontier, a learning principle that admits a world-model-guided decision only when its predicted advantage exceeds a certified bound on decision-relevant world-model error; otherwise, evidence is allocated to world-model verification. Action-conditioned value bounds and a closed-loop extension guarantee non-decreasing return for admitted decisions. Calibrated gates and simultaneous confidence sequences support adaptive evidence reuse, with sufficient and necessary verification bounds. Controlled learned-model experiments validate the predicted failure modes and certification behavior, while cross-backbone tool-use benchmarks instantiate the same verify-then-promote rule in realistic agent world-model pipelines, consistently improving decision quality and reliability.
Problem

Research questions and friction points this paper is trying to address.

world model
decision failure
failure attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Frontier
world model verification
certified bound
adaptive evidence reuse
decision quality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Huatai Zhu
Central South University
Qiang Chen
Qiang Chen
Professor of Informantion Engineering, Zhejiang University of Technology
Adaptive ControlFinite-time ControlNeural NetworksServo Systems
Z
Ziqian Kou
Xiangjiang Laboratory
W
Wenhao Li
University of Sydney
Fei Wang
Fei Wang
University of Science and Technology of China
CVDLVLAVLM
Y
Yichao Cao
Central South University
X
Xiu Su
Central South University
Yi Chen
Yi Chen
Hong Kong University of Science and Technology