Score
Designs and implements end‑to‑end pipelines that extract and align behavioral features from multiple data modalities (e.g., video, audio, sensors, logs), convert continuous signals into interpretable states, and produce structured behavioral representations; and analyzes those representations to discover cross‑modal associations, temporal patterns, and the stability or transferability of mined rules across contexts.
This work addresses the limitations of traditional product analytics, which rely on user-initiated queries and struggle to uncover unknown behavioral patterns due to high expertise barriers. The authors propose a behavior intelligence platform that transforms raw event streams into interpretable behavioral insights through a four-layer architecture, shifting the paradigm from passive response to proactive discovery. Key innovations include a formal definition of behavior intelligence, a taxonomy of phenomenon detectors, and an attention-constrained interestingness scoring mechanism. The system integrates semantic state normalization, absorbing Markov chain modeling of user journeys, and a large language model enhanced with behavioral knowledge graphs and factual constraints. This end-to-end framework autonomously identifies high-value behaviors and generates reliable narratives, substantially lowering the barrier to behavioral analysis and significantly enhancing the discovery of previously unknown patterns.
Existing models are largely restricted to unimodal or bimodal architectures, limiting their capacity to efficiently integrate the rich diversity of visual modalities present in real-world scenarios. To address this, we propose a low-supervision, fully automated multimodal video data pipeline that enables programmable composition and joint learning across heterogeneous visual modalities—including RGB, depth, optical flow, and edge maps. Our key contributions are: (1) PHG-MAE, a lightweight multimodal self-supervised encoder (<1M parameters), which leverages pretrained expert models and knowledge distillation to achieve performance on par with 300M-parameter large models; and (2) seamless integration of off-the-shelf modules (e.g., DPT) to enable real-time semantic segmentation and near-real-time depth estimation from handheld or webcam video on commodity hardware. Extensive experiments demonstrate the pipeline’s efficiency, scalability, and strong generalization under resource-constrained conditions.
This work addresses the challenge of detecting and regulating sycophantic behavior—excessive user flattery—in language models by proposing an iterative data generation method based on cascaded linear samples. Departing from conventional binary contrastive examples, the approach constructs sequences of samples with continuously varying behavioral intensities, revealing for the first time a linearly separable structure of sycophancy in activation space. This enables precise identification and disentanglement of the associated feature subspace. Through activation manipulation and subspace analysis, the method matches or exceeds baseline approaches such as LLM-as-a-judge and system prompting in detection accuracy, calibration, and robust controllability, while incurring lower computational overhead and substantially improving the interpretability of behavioral interventions.
Formal specification of embodied agent behaviors remains challenging due to the difficulty of rigorously encoding perception-driven, dynamic actions within traditional logical frameworks. Method: This paper introduces Embedding Temporal Logic (ETL), the first formal specification language that directly integrates semantic embeddings from pretrained vision and multimodal foundation models. ETL defines behavioral properties as distances between ideal behavior representations and actual observed representations in the embedding space, thereby overcoming limitations of classical logics in capturing perceptually grounded dynamics. The approach unifies embedding-guided specification, formal verification, robot planning, and foundation-model-based control. Contribution/Results: Evaluated on large-model-driven robotic tasks, ETL significantly enhances behavioral interpretability and controllability, enabling precise modeling and targeted steering of desired behaviors while preserving formal guarantees.
Behavioral modeling in robotics lacks systematic empirical understanding of the practical differences and commonalities between Behavior Trees (BTs) and State Machines (SMs). Method: We conduct the first large-scale empirical comparison across 1,200+ open-source ROS projects, leveraging domain-specific language (DSL) parsing, code mining, and conceptual mapping to analyze BT and SM usage across language design, structural abstraction, reuse patterns, and engineering practice. Contribution/Results: We find a significant upward trend in BT DSL adoption; uncover deep isomorphisms between BTs and SMs in control-flow abstraction granularity and modular reuse mechanisms; and release RoboBT-SM-Bench—the first cross-DSL, fully annotated benchmark dataset of robotic behavioral models. This work establishes an empirical foundation and infrastructure support for unifying theoretical frameworks and designing reusable architectures for behavioral modeling languages.
Current post-training of language models relies on abstract scalar rewards, which lack transparency regarding the instructional content of preference data and can lead models to learn spurious correlations, resulting in undesirable behaviors such as excessive stylization or sycophancy. This work proposes a data-centric post-training framework that, for the first time, leverages interpretability methods to explicitly model latent conceptual signals within preference data. By analyzing and identifying key features that distinguish preferred from non-preferred responses prior to optimization, the approach integrates interpretability protocols, statistical hypothesis testing, and fine-grained interventions at both feature and data levels. This enables effective diagnosis and suppression of harmful learning signals, significantly reducing off-target behaviors across multiple benchmarks while enhancing model safety and controllable personality.
本文提出了一种名为Behaviora的概念架构,通过行为片段、风格配置和体验配置来表示机器人与代理的内外部行为,并使用行为编译器将这些表示映射到具体平台动作。
This work addresses the challenge that existing vision-language-action models struggle to accurately execute natural language instructions involving spatiotemporal and logical constraints, while also lacking interpretability. The authors propose a hierarchical framework that, for the first time, deeply integrates Signal Temporal Logic (STL) between language understanding and robotic execution. The approach decomposes high-level instructions into subtasks and generates verifiable, optimizable, and correctable STL specifications, which dynamically schedule low-level policies. By combining vision-language models, STL, model predictive control, and learned policies, the method enables an end-to-end mapping from natural language instructions to formal specifications, supporting online monitoring and replanning. Experiments in real-world tabletop environments demonstrate significant improvements in accuracy, reliability, and interpretability of language-guided robotic tasks.
本文通过构建从模型到系统的分类法,系统化地解决了多模态学习中的计算、内存和部署瓶颈问题,并探讨了效率在不同层面的体现及优化方法。
该研究通过构建一个名为Ptolemy的语义地图来解决探索性数据分析过程中难以追踪已分析内容的问题,利用结构化描述生成嵌入式位置表示每个分析步骤,以提高全局定位和局部比较能力。