Beyond Appearance Shifts: Task-Semantic Action Calibration for VLA Models
本文提出BAS-VLA框架,通过任务语义动作校准解决VLA模型在任务语义变化下的行为稳定性与敏感性平衡问题。
本文提出BAS-VLA框架,通过任务语义动作校准解决VLA模型在任务语义变化下的行为稳定性与敏感性平衡问题。
本文提出了一种基于新冠病毒机制映射的冠状病毒优化算法(COA),用于解决全局优化问题,通过多种搜索操作如精英引导吸引、试向量生成等方法提高性能。
This work addresses the challenge of high-proportion missing data across arbitrary modality combinations in multimodal learning by proposing UL4M4, a task-agnostic, lightweight, and universally applicable unsupervised modality imputation framework. The method leverages modality-specific normalization and a novel partial-modality distance metric to enable fair clustering under frozen encoders, with cluster centroids guiding an iterative greedy imputation process. UL4M4 is the first approach to support decoupled imputation for any number of modalities and arbitrary missing patterns while effectively preserving cross-modal structural and scale invariance. Experimental results demonstrate that, even under extreme settings with over 50% missing modalities, UL4M4 consistently achieves F1-Micro scores above 0.7, significantly outperforming existing methods and exhibiting robustness across varying clustering scales.
This work addresses the performance degradation in multimodal imitation learning caused by missing visual or language inputs. The authors propose an end-to-end framework that operates without retraining, integrating a reinforcement learning–guided retrieval mechanism—based on Proximal Policy Optimization (PPO) and breadth-first search—to select the most relevant demonstrations from an expert dataset. Action signals are fused via soft cross-attention, and when modalities are missing, dedicated retrieval strategies combined with an embedding imputation head dynamically reconstruct the absent information. Experiments on three LIBERO benchmarks demonstrate that the proposed method significantly outperforms existing imitation learning approaches and maintains high robustness and performance even under sensor failure conditions.
This work addresses a critical security gap in large language model (LLM) agents that rely on user approval for sensitive actions: their self-generated summaries may be manipulated, leading users to authorize operations that differ from what is actually executed. To mitigate this risk, the paper introduces “consent integrity,” a novel security property formalizing the alignment between user-perceived intent and actual execution. Drawing inspiration from WYSIWYS (“What You See Is What You Sign”) and trusted path concepts, the authors propose a trusted intermediary situated at the agent–executor boundary. This intermediary enforces consent integrity through boundary event decoding, trusted rendering, and binding of displayed content to execution semantics. Evaluation on GTFOBins shows the prototype silently permits 10.0% of high-risk commands, while on tldr it flags 87.0% as non-reviewable, revealing a fundamental trade-off between security and usability.
本文提出BAS-VLA框架,通过任务语义动作校准解决VLA模型在任务语义变化下的行为稳定性与敏感性平衡问题。
本文提出了一种基于新冠病毒机制映射的冠状病毒优化算法(COA),用于解决全局优化问题,通过多种搜索操作如精英引导吸引、试向量生成等方法提高性能。
This work addresses the challenge of high-proportion missing data across arbitrary modality combinations in multimodal learning by proposing UL4M4, a task-agnostic, lightweight, and universally applicable unsupervised modality imputation framework. The method leverages modality-specific normalization and a novel partial-modality distance metric to enable fair clustering under frozen encoders, with cluster centroids guiding an iterative greedy imputation process. UL4M4 is the first approach to support decoupled imputation for any number of modalities and arbitrary missing patterns while effectively preserving cross-modal structural and scale invariance. Experimental results demonstrate that, even under extreme settings with over 50% missing modalities, UL4M4 consistently achieves F1-Micro scores above 0.7, significantly outperforming existing methods and exhibiting robustness across varying clustering scales.
This work addresses the performance degradation in multimodal imitation learning caused by missing visual or language inputs. The authors propose an end-to-end framework that operates without retraining, integrating a reinforcement learning–guided retrieval mechanism—based on Proximal Policy Optimization (PPO) and breadth-first search—to select the most relevant demonstrations from an expert dataset. Action signals are fused via soft cross-attention, and when modalities are missing, dedicated retrieval strategies combined with an embedding imputation head dynamically reconstruct the absent information. Experiments on three LIBERO benchmarks demonstrate that the proposed method significantly outperforms existing imitation learning approaches and maintains high robustness and performance even under sensor failure conditions.
This work addresses a critical security gap in large language model (LLM) agents that rely on user approval for sensitive actions: their self-generated summaries may be manipulated, leading users to authorize operations that differ from what is actually executed. To mitigate this risk, the paper introduces “consent integrity,” a novel security property formalizing the alignment between user-perceived intent and actual execution. Drawing inspiration from WYSIWYS (“What You See Is What You Sign”) and trusted path concepts, the authors propose a trusted intermediary situated at the agent–executor boundary. This intermediary enforces consent integrity through boundary event decoding, trusted rendering, and binding of displayed content to execution semantics. Evaluation on GTFOBins shows the prototype silently permits 10.0% of high-risk commands, while on tldr it flags 87.0% as non-reviewable, revealing a fundamental trade-off between security and usability.