WOLF: World Model Guided LiDAR Exploration with Predictive Frontiers
该研究提出WOLF框架,通过预测未来观测来改进基于LiDAR的无人机探索,使用循环世界模型学习观测动态,并结合预测前沿机制指导进一步感知。
该研究提出WOLF框架,通过预测未来观测来改进基于LiDAR的无人机探索,使用循环世界模型学习观测动态,并结合预测前沿机制指导进一步感知。
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
为解决机器人操作中动作表示与控制频率纠缠或依赖固定时间参数化的问题,提出CAT框架,通过编码动作轨迹为连续潜在标记,并加入频率感知位置编码,提高了不同控制频率下的表现。
This work addresses the limited dynamic depth selection in conventional Transformer residual architectures, which stems from insufficient historical information exchange across multiple parallel streams. To overcome this, we propose a reciprocal cross-stream addressing mechanism that enables bidirectional historical retrieval within a multi-stream framework: each stream computes depth weights based on the state of its counterpart stream and applies these weights to its own historical values. Our approach uniquely integrates inter-stream interaction into historical retrieval, preserving inter-layer representational diversity through cross-stream depth selection while mitigating redundancy and functional imbalance. Key components include reciprocal cross-attention, normalized state weighting, constrained gated writing, and block-level history storage. Experiments demonstrate consistent and significant improvements over standard residual Transformers and Attention Residuals across dense models (0.1B–1B) and a 7B sparse MoE model, with ablation studies confirming that performance gains arise from cross-stream interaction rather than additional parameters or projections.
This work addresses a key limitation in existing robotic manipulation approaches, which often neglect the dependence of action semantics on environmental context, resulting in noisy, redundant, and poorly structured control trajectories. To overcome this, the paper introduces EDAR—a novel framework that explicitly models the coupling between actions and their surrounding context. EDAR constructs environment-dependent action tokens by jointly embedding executable control commands with their visual outcomes in specific scenes. This representation enables the action space to capture interaction semantics rather than merely encoding command patterns. Evaluated in both simulated and real-world robotic manipulation tasks, EDAR significantly enhances downstream policy learning performance, demonstrating particularly strong gains in long-horizon manipulation scenarios.
该研究提出WOLF框架,通过预测未来观测来改进基于LiDAR的无人机探索,使用循环世界模型学习观测动态,并结合预测前沿机制指导进一步感知。
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
为解决机器人操作中动作表示与控制频率纠缠或依赖固定时间参数化的问题,提出CAT框架,通过编码动作轨迹为连续潜在标记,并加入频率感知位置编码,提高了不同控制频率下的表现。
This work addresses the limited dynamic depth selection in conventional Transformer residual architectures, which stems from insufficient historical information exchange across multiple parallel streams. To overcome this, we propose a reciprocal cross-stream addressing mechanism that enables bidirectional historical retrieval within a multi-stream framework: each stream computes depth weights based on the state of its counterpart stream and applies these weights to its own historical values. Our approach uniquely integrates inter-stream interaction into historical retrieval, preserving inter-layer representational diversity through cross-stream depth selection while mitigating redundancy and functional imbalance. Key components include reciprocal cross-attention, normalized state weighting, constrained gated writing, and block-level history storage. Experiments demonstrate consistent and significant improvements over standard residual Transformers and Attention Residuals across dense models (0.1B–1B) and a 7B sparse MoE model, with ablation studies confirming that performance gains arise from cross-stream interaction rather than additional parameters or projections.
This work addresses a key limitation in existing robotic manipulation approaches, which often neglect the dependence of action semantics on environmental context, resulting in noisy, redundant, and poorly structured control trajectories. To overcome this, the paper introduces EDAR—a novel framework that explicitly models the coupling between actions and their surrounding context. EDAR constructs environment-dependent action tokens by jointly embedding executable control commands with their visual outcomes in specific scenes. This representation enables the action space to capture interaction semantics rather than merely encoding command patterns. Evaluated in both simulated and real-world robotic manipulation tasks, EDAR significantly enhances downstream policy learning performance, demonstrating particularly strong gains in long-horizon manipulation scenarios.