MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of detecting and correcting unsupervised, silent, and cumulative latent hallucinations in world models. To this end, it proposes MEND, a framework that pioneers the use of a single conditional score network to unify the detection, localization, and inference-time correction of hallucinations without labeled data. By leveraging masked empirical Bayes neural denoising and denoising score matching, the method constructs a unified score field without requiring ground-truth error labels. Experimental results on navigation tasks demonstrate that MEND achieves an AUROC of 0.80, effectively suppressing the accumulation of single-step latent errors and significantly improving predictive performance.
📝 Abstract
World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.
Problem

Research questions and friction points this paper is trying to address.

World Models
Latent Hallucination
Error Detection
Error Localisation
Self-supervised Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Latent Hallucination
Score Matching
Masked Empirical-Bayes Neural Denoising
Label-Free Detection
🔎 Similar Papers
No similar papers found.