Score
Designs and implements encoding schemes that map observed behaviors into paired or dual identifiers (for example a categorical behavior code coupled with a textual descriptor) and/or transform sequences of behavioral events into a single token stream. Builds annotation schemas, tokenizers, and converters that preserve temporal and semantic structure so behaviors can be integrated into model inputs and downstream computational analyses.
Current Facial Action Coding Systems (FACS) face two key bottlenecks in mental health research: limited accuracy in Action Unit (AU) detection and incomplete coverage of facial movements, hindering comprehensive facial expression representation. This paper introduces Facial Basis—the first unsupervised, data-driven 3D facial motion coding framework—replacing hand-crafted AUs with interpretable, localized motion primitives to enable complete, additive modeling. Its core contributions are: (1) constructing the first unsupervised facial motion dictionary; (2) overcoming FACS’s three fundamental limitations—reliance on manual annotation, incomplete coverage, and non-additive primitives; and (3) establishing an end-to-end AU-based comparative evaluation framework. On real-world conversational videos, Facial Basis achieves significantly higher accuracy than state-of-the-art AU detectors in predicting autism spectrum disorder diagnosis. The code and models are publicly released.
Existing interpretable recommendation methods rely on ID-based representations, which suffer from semantic ambiguity and poor compatibility with large language models (LLMs); moreover, user behaviors exhibit entangled multi-intent patterns, and collaborative signals are semantically misaligned with natural language. To address these issues, we propose a **behavior tokenization paradigm**: leveraging graph neural networks to learn structured user–item interaction representations, and employing vector-quantized variational autoencoders (VQ-VAEs) to disentangle macro-level interests from micro-level intentions, thereby constructing a transferable, graph-enhanced behavior lexicon. We further design a multi-level semantic supervision scheme and an LLM input embedding alignment mechanism—freezing LLM embeddings—to bridge behavioral signals with natural language semantics. Evaluated on three public benchmarks, our method significantly improves zero-shot recommendation performance, generates coherent and informative explanations, and yields behavior tokens with fine-grained interpretability and cross-domain transferability.
To address inefficiencies in action sequence encoding—such as low tokenization efficiency for discrete/continuous representations, non-smooth trajectories, reliance on auxiliary training modules, and poor compatibility with large models—this paper proposes BEAST, a B-spline-based action tokenizer. BEAST introduces a novel, training-free, fixed-length tokenization paradigm: it directly maps continuous actions to compact discrete or continuous tokens via B-spline parameterization, inherently ensuring trajectory smoothness and high-frequency control capability. The method seamlessly integrates with VAEs, Transformers, and multimodal large models (e.g., Florence-2), supporting parallel decoding and end-to-end training. Evaluated on 166 simulated and 8 real-robot tasks, BEAST significantly reduces training and inference overhead while generating high-fidelity, smooth control signals, achieving state-of-the-art task success rates.
Existing models for smart home behavior analysis, trained on static datasets, exhibit poor generalization under behavioral drift caused by seasonal changes and evolving user habits; re-collecting and annotating new data is costly, time-consuming, and raises privacy concerns. Method: We propose SmartGen—a novel framework integrating time- and semantic-aware subsequence partitioning, latent-space behavioral clustering compression, graph-guided sequence generation, and a two-stage anomaly filtering mechanism—to leverage large language models (LLMs) for synthesizing high-fidelity, context-aware user behavior sequences that enable continual downstream model adaptation. Contribution/Results: Evaluated on three real-world smart home datasets, SmartGen achieves average improvements of 85.43% in anomaly detection and 70.51% in behavior prediction over state-of-the-art methods, demonstrating superior robustness to behavioral drift without requiring new labeled data.
Lengthy system prompts in large language models (LLMs) incur high inference latency, substantial computational overhead, and inefficient context-space utilization. Method: We propose a single-token prompt compression method that requires neither internal model access nor annotated data. Our approach employs a lightweight three-stage training framework integrating reconstruction-based content encoding and behavioral distillation to self-supervise the semantic and functional condensation of system prompts into a single learnable token. Contribution/Results: Evaluated on three benchmark datasets, our method achieves up to 3000× prompt-length compression, significantly reducing inference cost while retaining ~98% of original task performance. To the best of our knowledge, this is the first work achieving high-fidelity, single-token system prompt compression under zero-shot, black-box conditions—establishing a novel paradigm for efficient LLM system prompt deployment.
Behavioral profiling (BP) annotation is challenging to automate due to its multidimensional, multilingual nature, and conventional task-level evaluation obscures underlying skill heterogeneity. This work proposes a novel “skill feasibility” paradigm, decomposing BP annotation into 14 operationalizable annotation skills and implementing a schema-guided, skill-document-driven pipeline. Evaluation over a 300-instance validation set—through two rounds of testing involving human annotators and large language models (GPT-5.4 and three open-source models)—reveals a “shared categorization, independent execution” pattern: humans and GPT exhibit high agreement at the skill level but diverge in instance-level execution. The study identifies five directly feasible skills, four recoverable via relabeling, and five structurally undefined. GPT-5.4 demonstrates reliable performance on feasible skills (accuracy = 0.678, κ = 0.665, weighted F1 = 0.695), whereas open-source models primarily fail in translating schemas into executable skills.
This study addresses the lack of structured methodologies for behavior development on resource-constrained robotic platforms by conducting a systematic analysis of the R-CODE script corpus from Sony’s ERS-111 AIBO robot. The work proposes state-level abstraction as an intermediate representation and identifies core behavioral primitives—including initialization, perception, iterative action, synchronization, and recovery. Through behavior graph analysis, cross-script comparison, and finite state machine modeling, the authors distill a compact embodied behavior grammar, demonstrating that a wide range of robot behaviors can be composed from a small set of generic control structures. This contribution establishes a reusable, structured foundation for modular behavior synthesis in deterministic, hardware-direct control scenarios.
This work investigates the mechanistic nature of behavioral changes in language model post-training—whether such changes arise from rewriting, creating new mechanisms, or merely reweighting existing ones. To this end, the authors propose Behavior Manifold Analysis, a method that constructs low-dimensional local charts in both activation space (ACT) and neuron output contribution space (NOC) to trace how supervised fine-tuning (SFT) and reward optimization reshape behavioral geometry. For the first time from a geometric perspective, the study reveals that SFT substantially reconstructs the behavior manifold, whereas reward optimization largely preserves its underlying structure and primarily adjusts output weights. Experiments across multiple model architectures demonstrate the high compressibility of these charts and their partial cross-model alignment, highlighting the method’s novel utility for mechanistic interpretability and cross-architecture comparison.
This work addresses the challenge of efficiently converting graph-structured data into sequences compatible with general-purpose Transformer models. The authors propose a novel graph tokenization framework that, for the first time, integrates reversible graph serialization—guided by global substructure frequency statistics—with Byte Pair Encoding (BPE) to produce compact token representations that preserve structural semantics while remaining amenable to sequence-based architectures. Notably, this approach requires no modifications to standard Transformer backbones such as BERT, enabling direct application to graph data. Evaluated across 14 established graph benchmark datasets, the method significantly outperforms both conventional graph neural networks and specialized graph Transformers, achieving state-of-the-art performance.