interaction log analysis

Designs and implements pipelines and methods to extract, preprocess, and analyze recorded interaction events (usage logs, traces) to produce behavioral features, detect and categorize anomalous or misaligned actions, and compute metrics derived from event sequences. Builds statistical and probabilistic models to estimate event or sequence log-probabilities and their shifts, measure the impact of interventions on retrieval or outcomes, and correlate log-derived signals with user- or system-level differences.

interactionloganalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$194K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.

analytical workflowcomputer-based assessmentsconsistency check

Accurate and Noise-Tolerant Extraction of Routine Logs in Robotic Process Automation (Extended Version)

Oct 09, 2025
MD
Massimiliano de Leoni
🏛️ University of Padua | Universitas Mercatorum of Rome

Existing work primarily focuses on action-set extraction, neglecting end-to-end routine model discovery and lacking validation on real-world UI logs corrupted by execution variability and human errors. This paper proposes a noise-tolerant clustering method that, for the first time, directly enables complete and high-precision extraction of routine logs from raw UI interaction traces—thereby facilitating routine pattern discovery in robotic process automation (RPA). Our approach integrates behavioral similarity measurement with an adaptive noise-filtering mechanism to robustly identify and reconstruct routine execution paths. Extensive experiments across nine publicly available UI log datasets demonstrate that our method achieves an average F1-score improvement of over 15% under high-noise conditions, significantly outperforming state-of-the-art techniques. The key contributions include: (i) the first end-to-end routine discovery framework tailored for noisy UI logs; (ii) a principled noise-resilient clustering strategy grounded in behavioral semantics; and (iii) empirically validated superiority in both accuracy and robustness.

Enabling robotic process automation through routine-type model discoveryExtracting accurate routine logs from user interface interactionsHandling inconsistent routine execution and noise in process data

A Ground Truth Approach for Assessing Process Mining Techniques

Jan 24, 2025
DS
Dominique Sommers
🏛️ Eindhoven University of Technology

Existing process mining evaluation faces challenges including the scarcity of real-world logs with ground-truth models, noisy event logs, and a disconnect between behavioral deviation modeling and log generation. This paper proposes the first traceable synthetic benchmark framework jointly linking models, logs, and deviations: taking an initial Petri net or BPMN process model as input, it integrates a library of behavioral deviation patterns (e.g., skipping, redoing, reordering) and event-level log perturbation mechanisms to generate imperfect logs with controllable deviations. Unlike conventional approaches injecting noise only at the log level, our framework enables fine-grained, joint modeling of both behavioral deviations and recording errors. Evaluation is systematically performed via relaxed alignments. We construct three synthetic datasets; one has been successfully deployed in conformance checking evaluation, quantitatively exposing algorithmic sensitivity differences and qualitatively characterizing their explanatory boundaries across distinct deviation types.

Accuracy ImprovementData SimulationWorkflow Analysis

Performance Analysis: Discovering Semi-Markov Models From Event Logs

Jun 29, 2022
AK
A. Kalenkova
🏛️ The University of Adelaide | Adelaide Data Science Centre

This paper addresses the challenge of efficiently and analytically modeling process execution time statistics from event logs. Methodologically, it introduces the first end-to-end analytical performance analysis framework based on semi-Markov processes: it directly infers execution time means and probability density functions (PDFs) from logs—bypassing simulation entirely. For discrete-time execution times, it employs exact convolution; for continuous-time cases, it approximates PDFs using Gaussian mixture models (GMMs), balancing accuracy, model compactness, and interpretability. Experiments show that the discrete-time approach achieves up to one order of magnitude speedup over simulation under small support sets, while GMM-based representation drastically reduces model size, enabling rapid what-if analysis. The core contribution is the first fully analytical, log-driven inference of semi-Markov performance models—eliminating reliance on traditional simulation-based approaches and establishing a new paradigm for scalable, interpretable process performance analysis.

Develops analytical techniques for performance analysis using semi-Markov processes.Estimates mean execution time and builds probability density functions for process execution.Provides efficient, simulation-free solutions for what-if analysis in process mining.

LogLM: From Task-based to Instruction-based Automated Log Analysis

Oct 12, 2024
YL
Yilun Liu
🏛️ Huawei | Nankai University

Existing log analysis models are task-specific, rely heavily on domain-specific annotated data, exhibit poor generalization, and struggle with complex or unseen instructions. Method: We propose LogLM, an instruction-driven large language model for log analysis, which unifies diverse log tasks—including anomaly detection, parsing, and summarization—into a standardized instruction-response format. LogLM is adapted to the log domain via multi-task instruction tuning and log-specific instruction engineering. It accepts natural-language instructions and supports zero-shot cross-task transfer. Contribution/Results: Experiments demonstrate that LogLM outperforms all state-of-the-art methods across five core log analysis tasks. It exhibits strong generalization to complex instructions and previously unseen tasks. As a single unified model, LogLM replaces multiple specialized models, significantly improving deployment efficiency and task-agnostic capability.

Log AnalysisModel AdaptabilityTask Generalization

Latest Papers

What's happening recently
View more

This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.

downstream tasksirrelevant messageslog analysis

This work addresses the challenge of inefficient anomaly diagnosis due to the unstructured and semantically impoverished nature of traditional system logs. The authors propose a hierarchical log abstraction method that parses raw logs into a three-layer semantic structure—entities, actions, and states—and introduce a modular collaborative detection framework that performs anomaly detection at each semantic level. By integrating large language models (LLMs) with a human-in-the-loop interactive visualization system, the approach enables precise identification, localization, and interpretable analysis of anomalies. Evaluated on the HDFS benchmark dataset, the method demonstrates effectiveness while supporting hierarchical log browsing, highlighting of anomalous segments, and user-guided review and correction of LLM-generated explanations. The source code and an online demo platform have been publicly released.

anomaly diagnosisexecution behavior understandinghierarchical log abstraction

This work addresses the challenges of efficiently analyzing large-scale, dynamically evolving semi-structured logs under conditions of label scarcity and distribution shift, which hinder system reliability and AIOps advancement. It presents the first unified task taxonomy for log analysis driven by large language models (LLMs), offering a systematic survey of their application across the full log analysis pipeline—including log generation, parsing, anomaly detection, and root cause analysis. Through structured analysis of 145 studies, the paper identifies five core design paradigms: prompt engineering, retrieval augmentation, fine-tuning, agent collaboration, and result verification. It further synthesizes the state of research, datasets, and evaluation practices across seven key tasks, while highlighting critical challenges in robustness, trustworthiness, and reproducibility, thereby providing a comprehensive roadmap for reliable LLM-based log intelligence.

AIOpsdata driftlarge language models

Low-level user interaction logs are often noisy and fine-grained, making it challenging to extract interpretable, high-level behavioral patterns across applications. This work proposes WorkflowView, the first large language model (LLM)-based framework for cross-domain abstraction of action sequences, which maps raw interaction logs to high-level semantic workflows through semantic similarity computation and few-shot learning. The approach demonstrates strong generalization and inherent privacy-preserving properties in both zero-shot and few-shot settings: it achieves a semantic similarity of 0.91 in browser log task reconstruction and a weighted F1 score of 0.90 in MOOC dropout prediction. Furthermore, WorkflowView has been successfully applied to anonymized analysis of AI tool usage within Microsoft Word, highlighting its practical utility in real-world scenarios.

action sequence abstractionbehavioral data analysiscross-domain generalization

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

Hot Scholars

JL

Jionghao Lin

University of Hong Kong | Carnegie Mellon University | Monash University
Artificial Intelligence in EducationLearning AnalyticsHuman-Centered AIFeedback
JP

James Prather

Associate Professor of Computer Science, Abilene Christian University
Human-Computer InteractionComputer Science EducationGenerative AI
JL

Juho Leinonen

Aalto University
Computing EducationLearning AnalyticsGenerative AIAI in Education
PD

Paul Denny

Professor, University of Auckland
Educational technologyComputer Science Education
EC

Eason Chen

Human-Computer Interaction Institute, Carnegie Mellon University
Learning SciencesEducation TechnologiesLearning AnalyticsBlockchain