Score
Designs, builds, and analyzes models and analytics that capture, represent, and predict sequences and patterns of entity behavior, including feature extraction, sequence and tree/graph representations, clustering and segmentation, pattern mining, and behavior-prediction models. Develops and applies tests, probing methods, and quantitative evaluation metrics to validate, compare, and interpret behavioral models and mined behavioral patterns.
This work addresses the limitations of traditional product analytics, which rely on user-initiated queries and struggle to uncover unknown behavioral patterns due to high expertise barriers. The authors propose a behavior intelligence platform that transforms raw event streams into interpretable behavioral insights through a four-layer architecture, shifting the paradigm from passive response to proactive discovery. Key innovations include a formal definition of behavior intelligence, a taxonomy of phenomenon detectors, and an attention-constrained interestingness scoring mechanism. The system integrates semantic state normalization, absorbing Markov chain modeling of user journeys, and a large language model enhanced with behavioral knowledge graphs and factual constraints. This end-to-end framework autonomously identifies high-value behaviors and generates reliable narratives, substantially lowering the barrier to behavioral analysis and significantly enhancing the discovery of previously unknown patterns.
A systematic literature review on the application of large language models (LLMs) to behavioral modeling—particularly automated generation of use case and sequence diagrams—is currently lacking, hindering research consolidation and practical guidance. Method: This paper presents the first comprehensive survey in this domain, identifying 14 core studies via a terminology-driven search strategy and synthesizing prevalent LLM application patterns and evaluation methodologies for behavioral modeling. Results: Findings confirm the feasibility of LLMs for diagram generation tasks; however, existing work is heavily reliant on GPT-series models and largely omits validation by domain experts. The study innovatively advocates for cross-model comparative analysis and expert-in-the-loop evaluation. It thereby provides theoretical foundations and actionable pathways for future research, tool development, and pedagogical practice in model-driven engineering and AI-assisted software modeling.
The absence of automated, systematic approaches for modeling pattern recognition in conceptual modeling hinders improvements in knowledge representation and modeling quality. Method: This paper introduces frequent subgraph mining—systematically applied to conceptual modeling for the first time—and proposes a cross-language (OntoUML/ArchiMate), multi-criteria structural pattern discovery framework. It integrates a gSpan variant with graph editing, graph isomorphism testing, and pattern abstraction techniques to build an extensible, exploratory analysis tool. Contribution/Results: Evaluated on two authoritative datasets, the method successfully identifies highly reusable structural patterns. It demonstrates effectiveness in assessing modeling practices, supporting language evolution, and optimizing model quality—thereby filling a critical research gap in automated pattern mining for conceptual modeling.
Behavioral modeling in robotics lacks systematic empirical understanding of the practical differences and commonalities between Behavior Trees (BTs) and State Machines (SMs). Method: We conduct the first large-scale empirical comparison across 1,200+ open-source ROS projects, leveraging domain-specific language (DSL) parsing, code mining, and conceptual mapping to analyze BT and SM usage across language design, structural abstraction, reuse patterns, and engineering practice. Contribution/Results: We find a significant upward trend in BT DSL adoption; uncover deep isomorphisms between BTs and SMs in control-flow abstraction granularity and modular reuse mechanisms; and release RoboBT-SM-Bench—the first cross-DSL, fully annotated benchmark dataset of robotic behavioral models. This work establishes an empirical foundation and infrastructure support for unifying theoretical frameworks and designing reusable architectures for behavioral modeling languages.
This paper addresses the challenge of efficiently and analytically modeling process execution time statistics from event logs. Methodologically, it introduces the first end-to-end analytical performance analysis framework based on semi-Markov processes: it directly infers execution time means and probability density functions (PDFs) from logs—bypassing simulation entirely. For discrete-time execution times, it employs exact convolution; for continuous-time cases, it approximates PDFs using Gaussian mixture models (GMMs), balancing accuracy, model compactness, and interpretability. Experiments show that the discrete-time approach achieves up to one order of magnitude speedup over simulation under small support sets, while GMM-based representation drastically reduces model size, enabling rapid what-if analysis. The core contribution is the first fully analytical, log-driven inference of semi-Markov performance models—eliminating reliance on traditional simulation-based approaches and establishing a new paradigm for scalable, interpretable process performance analysis.
This study addresses the challenge of automatically constructing high-fidelity user journey models from interaction logs between users and digital services. The authors propose a novel hybrid approach that integrates automata learning with process mining techniques. A key innovation is the introduction of an adaptive algorithm selection mechanism, which dynamically chooses the optimal modeling strategy based on the characteristics of the event log. This mechanism effectively mitigates the limitations of each individual technique—namely, their reliance on expert knowledge or assumptions about specific event distributions. Empirical evaluation on real-world datasets demonstrates that the proposed hybrid method significantly outperforms either technique used in isolation, yielding substantially more accurate user journey models.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.
AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess whether an evaluation worked as intended. Researchers have started developing methods for log analysis, but a standardised approach is still missing. Here we suggest a pipeline based on current best practices. We illustrate it with concrete code examples in the Inspect Scout library, provide detailed guidance on each step, and highlight common pitfalls. Our framework provides researchers with a foundation for rigorous and reproducible log analysis.