Institution profile

BYD Automotive Co., Ltd.

Industry researchasia · cn
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation

Oct 02, 2026

This study addresses the reliance of Vision-Language-Action (VLA) models on costly real-world data and their limited generalization in contact-rich manipulation tasks by proposing the SARI framework. Motivated by the insight that spatial coverage can be simulated while physical contacts require real-world grounding, SARI decouples tasks into two phases: simulated approach and real interaction. It leverages digital twins to generate diverse spatial trajectories and trains a unified policy with minimal real contact data. Seamless transfer without explicit labels is achieved through visual appearance alignment, shared camera-relative action representations, and co-training. Experimental results demonstrate that this approach reduces data collection time by 34.3% and achieves a 27.5% success rate on unseen object poses, significantly outperforming baselines.

0 citationsRead paper

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

Sep 28, 2026

This study addresses a tokenization blind spot in chat-template prompt injection, wherein reserved tokens and their synonymous subword-decoded counterparts exhibit identical semantics yet vastly different attack efficacy—a vulnerability overlooked by existing defenses. Leveraging the InjecAgent and AgentDojo benchmarks, this work employs controlled experiments to quantify embedding vectors, revealing that instruction authority stems from a single learned vector rather than subword averages, and demonstrating that instruction tuning amplifies the model’s preference for reserved tokens. Consequently, the proposed approach reduces attack success rates across most open-source models by 39 to 66 percentage points. Furthermore, this research identifies tool-protocol token vulnerabilities in 255 widely used models and establishes tokenization mechanisms as a critical decision point for security defenses.

0 citationsRead paper

One Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent Prediction

Sep 26, 2026

This study addresses the limitation of existing autonomous driving traffic signal perception, which merely recognizes colors without semantically interpreting right-of-way for specific driving maneuvers. To bridge this gap, this work proposes a directional traffic signal understanding task that predicts structured signal states and passing permissions—such as going straight or turning—directly from images. Accordingly, a direction-level benchmark is constructed based on OpenLane-V2, alongside a direction-aware baseline model designed to integrate global context, local evidence, and action-specific representations. By closing the semantic gap between perception and planning, the proposed approach provides an interpretable signal interface. Experimental results demonstrate that it significantly improves passing permission prediction accuracy in complex intersection scenarios compared to conventional detection pipelines, thereby effectively supporting downstream path planning.

0 citationsRead paper
Recent publications

Latest Papers

SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation

Oct 02, 2026

This study addresses the reliance of Vision-Language-Action (VLA) models on costly real-world data and their limited generalization in contact-rich manipulation tasks by proposing the SARI framework. Motivated by the insight that spatial coverage can be simulated while physical contacts require real-world grounding, SARI decouples tasks into two phases: simulated approach and real interaction. It leverages digital twins to generate diverse spatial trajectories and trains a unified policy with minimal real contact data. Seamless transfer without explicit labels is achieved through visual appearance alignment, shared camera-relative action representations, and co-training. Experimental results demonstrate that this approach reduces data collection time by 34.3% and achieves a 27.5% success rate on unseen object poses, significantly outperforming baselines.

0 citationsRead paper

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

Sep 28, 2026

This study addresses a tokenization blind spot in chat-template prompt injection, wherein reserved tokens and their synonymous subword-decoded counterparts exhibit identical semantics yet vastly different attack efficacy—a vulnerability overlooked by existing defenses. Leveraging the InjecAgent and AgentDojo benchmarks, this work employs controlled experiments to quantify embedding vectors, revealing that instruction authority stems from a single learned vector rather than subword averages, and demonstrating that instruction tuning amplifies the model’s preference for reserved tokens. Consequently, the proposed approach reduces attack success rates across most open-source models by 39 to 66 percentage points. Furthermore, this research identifies tool-protocol token vulnerabilities in 255 widely used models and establishes tokenization mechanisms as a critical decision point for security defenses.

0 citationsRead paper

One Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent Prediction

Sep 26, 2026

This study addresses the limitation of existing autonomous driving traffic signal perception, which merely recognizes colors without semantically interpreting right-of-way for specific driving maneuvers. To bridge this gap, this work proposes a directional traffic signal understanding task that predicts structured signal states and passing permissions—such as going straight or turning—directly from images. Accordingly, a direction-level benchmark is constructed based on OpenLane-V2, alongside a direction-aware baseline model designed to integrate global context, local evidence, and action-specific representations. By closing the semantic gap between perception and planning, the proposed approach provides an interpretable signal interface. Experimental results demonstrate that it significantly improves passing permission prediction accuracy in complex intersection scenarios compared to conventional detection pipelines, thereby effectively supporting downstream path planning.

0 citationsRead paper