Score
Designs and implements instrumentation and data pipelines to log, store, and process behavioral events and behavior-tree representations, and engineers scalable analytics for querying and aggregating those logs. Builds models and analyses of user behavior — including sequence- and tree-based models, segmentation, and metric computation — to characterize interaction patterns, detect anomalies, and evaluate behavior-driven hypotheses.
Business process optimization remains challenging due to fragmented methodologies across process mining, predictive process monitoring, and process-aware recommendation—each operating in isolation without a unified theoretical foundation or integration framework. Method: This paper proposes a closed-loop optimization framework that systematically integrates Alpha algorithm/Inductive Miner for process discovery, LSTM/Transformer for runtime prediction, collaborative filtering/graph neural networks for action recommendation, and explainable AI (XAI) for interpretability—enabling automated bottleneck identification, anomaly forecasting, and prescriptive optimization from event logs. Contribution/Results: We establish the first unified conceptual boundary, evolutionary taxonomy, and synergy paradigm across the three domains; construct a comprehensive classification schema covering 120+ studies; clarify application scopes and standardized evaluation benchmarks; and deliver an industrially actionable methodology selection guide with validated deployment pathways.
This work addresses the limitations of traditional product analytics, which rely on user-initiated queries and struggle to uncover unknown behavioral patterns due to high expertise barriers. The authors propose a behavior intelligence platform that transforms raw event streams into interpretable behavioral insights through a four-layer architecture, shifting the paradigm from passive response to proactive discovery. Key innovations include a formal definition of behavior intelligence, a taxonomy of phenomenon detectors, and an attention-constrained interestingness scoring mechanism. The system integrates semantic state normalization, absorbing Markov chain modeling of user journeys, and a large language model enhanced with behavioral knowledge graphs and factual constraints. This end-to-end framework autonomously identifies high-value behaviors and generates reliable narratives, substantially lowering the barrier to behavioral analysis and significantly enhancing the discovery of previously unknown patterns.
AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess whether an evaluation worked as intended. Researchers have started developing methods for log analysis, but a standardised approach is still missing. Here we suggest a pipeline based on current best practices. We illustrate it with concrete code examples in the Inspect Scout library, provide detailed guidance on each step, and highlight common pitfalls. Our framework provides researchers with a foundation for rigorous and reproducible log analysis.
Behavioral modeling in robotics lacks systematic empirical understanding of the practical differences and commonalities between Behavior Trees (BTs) and State Machines (SMs). Method: We conduct the first large-scale empirical comparison across 1,200+ open-source ROS projects, leveraging domain-specific language (DSL) parsing, code mining, and conceptual mapping to analyze BT and SM usage across language design, structural abstraction, reuse patterns, and engineering practice. Contribution/Results: We find a significant upward trend in BT DSL adoption; uncover deep isomorphisms between BTs and SMs in control-flow abstraction granularity and modular reuse mechanisms; and release RoboBT-SM-Bench—the first cross-DSL, fully annotated benchmark dataset of robotic behavioral models. This work establishes an empirical foundation and infrastructure support for unifying theoretical frameworks and designing reusable architectures for behavioral modeling languages.
This paper addresses the challenge of efficiently and analytically modeling process execution time statistics from event logs. Methodologically, it introduces the first end-to-end analytical performance analysis framework based on semi-Markov processes: it directly infers execution time means and probability density functions (PDFs) from logs—bypassing simulation entirely. For discrete-time execution times, it employs exact convolution; for continuous-time cases, it approximates PDFs using Gaussian mixture models (GMMs), balancing accuracy, model compactness, and interpretability. Experiments show that the discrete-time approach achieves up to one order of magnitude speedup over simulation under small support sets, while GMM-based representation drastically reduces model size, enabling rapid what-if analysis. The core contribution is the first fully analytical, log-driven inference of semi-Markov performance models—eliminating reliance on traditional simulation-based approaches and establishing a new paradigm for scalable, interpretable process performance analysis.
Existing log analysis models are task-specific, rely heavily on domain-specific annotated data, exhibit poor generalization, and struggle with complex or unseen instructions. Method: We propose LogLM, an instruction-driven large language model for log analysis, which unifies diverse log tasks—including anomaly detection, parsing, and summarization—into a standardized instruction-response format. LogLM is adapted to the log domain via multi-task instruction tuning and log-specific instruction engineering. It accepts natural-language instructions and supports zero-shot cross-task transfer. Contribution/Results: Experiments demonstrate that LogLM outperforms all state-of-the-art methods across five core log analysis tasks. It exhibits strong generalization to complex instructions and previously unseen tasks. As a single unified model, LogLM replaces multiple specialized models, significantly improving deployment efficiency and task-agnostic capability.
This work proposes an automated log aggregation and analysis framework based on large language models to address the growing challenge of log analysis in increasingly complex systems, where engineers traditionally rely on domain expertise to manually craft intricate LogQL queries. The framework enables end-to-end generation of LogQL queries from natural language instructions by integrating a hierarchical log knowledge base, natural language understanding, knowledge retrieval, and tool invocation mechanisms. Evaluated on four real-world log datasets, the approach achieves an average accuracy of 76.8%, significantly outperforming existing baselines and demonstrating its effectiveness and practicality for log analysis tasks.
This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.
Low-level user interaction logs are often noisy and fine-grained, making it challenging to extract interpretable, high-level behavioral patterns across applications. This work proposes WorkflowView, the first large language model (LLM)-based framework for cross-domain abstraction of action sequences, which maps raw interaction logs to high-level semantic workflows through semantic similarity computation and few-shot learning. The approach demonstrates strong generalization and inherent privacy-preserving properties in both zero-shot and few-shot settings: it achieves a semantic similarity of 0.91 in browser log task reconstruction and a weighted F1 score of 0.90 in MOOC dropout prediction. Furthermore, WorkflowView has been successfully applied to anonymized analysis of AI tool usage within Microsoft Word, highlighting its practical utility in real-world scenarios.