Score
Designs and implements systems that aggregate continuous personal sensor and activity streams and segment them into temporally localized events or episodes, each with a timestamp. Builds models and pipelines for event boundary detection, multimodal fusion, semantic labeling (personicle/semantic event detection), and contextualization of events with surrounding behavioral cues.
Existing evaluation approaches for streaming process mining algorithms predominantly rely on static logs or synthetic event streams, which fail to capture the complexity of real-world event streams in IoT environments—such as out-of-order events, concurrency, incomplete cases, and concept drift. This work addresses this gap by introducing, for the first time, a feature framework from data stream research into streaming process mining. It proposes an intent-oriented event stream generation methodology, extends the conceptual model of event streams, and implements a prototype tool, Stream of Intent. This tool enables customizable configuration of key stream characteristics reflective of real-world scenarios, facilitating the generation of controlled, reproducible, and realistically complex event streams. Consequently, it significantly enhances the relevance and adaptability of algorithm evaluation and development in streaming process mining.
This paper addresses the fragmentation between wearable-device data and process mining. We propose an event log enhancement method for personal behavior modeling, integrating smartwatch data (sleep, heart rate, physical activity) and digital calendar data into process event logs via temporal alignment, multi-source data aggregation, and event derivation. Novel semantic events—such as “deep sleep onset” and “prolonged sitting before meeting”—are introduced at the event, case, and activity levels. Three wearable-data fusion pathways are innovatively designed, relaxing process mining’s traditional reliance on structured business logs. Evaluated on 30 days of real-world, multimodal data from individual users, our approach significantly improves the granularity and interpretability of behavioral pattern discovery. It establishes a scalable, process-mining-based paradigm for personalized productivity optimization and holistic well-being analysis.
This work addresses the lack of a unified, reproducible evaluation platform for human activity recognition (AR) using binary sensor data. Methodologically, we propose the first modular end-to-end AR pipeline, comprising a data-driven three-stage framework: robust data cleaning, sliding-window-based adaptive temporal segmentation, and lightweight personalized classification (using XGBoost or LSTM variants), enabling plug-and-play substitution of methods, datasets, and evaluation protocols. Our key contribution is the first modular AR architecture specifically designed for binary sensing—ensuring full pipeline reproducibility, customization, and rigorous evaluation. Extensive validation across multiple public datasets demonstrates significant improvements in cross-user accuracy and generalization robustness. The platform establishes a standardized experimental baseline for AR research and supports rapid prototyping and deployment.
Existing user event modeling approaches treat individual behavioral sequences and relational interactions (e.g., social graphs) separately, failing to capture their synergistic effects in a unified framework. Method: We introduce the first publicly available dataset and standardized prediction task supporting *joint modeling* of individual and relational events. We propose a unified formalism that integrates temporal sequence modeling with graph neural networks to jointly represent dynamic behavioral sequences and static/dynamic social structures. Contribution/Results: Experiments demonstrate substantial improvements in event prediction performance: our method outperforms unimodal baselines by an average of 12.7% across multiple metrics. To foster reproducibility and community advancement, we release the dataset, source code, and benchmark tasks—establishing a new paradigm and foundational infrastructure for modeling complex, socially embedded user behavior.
This study addresses the challenge of transforming multimodal passive tracking data (from smartphones and wearables) into high-level, context-aware insights. We propose an LLM-augmented, human-in-the-loop visual analytics paradigm that integrates interactive visualization, lightweight large language model agents, and human-centered design to construct an expert-perception modeling framework. Through a three-round, technology-probe-based empirical study with 21 domain experts, we validate the system’s effectiveness in enhancing insight exploration efficiency and analytical credibility. Our contributions are threefold: (1) the first AI–visualization co-reasoning framework tailored for expert interpretation of personal tracking data; (2) seven reusable design principles for AI-enhanced visualization; and (3) open-sourced, reproducible expert-perception models and design insights to advance human–AI collaborative analytics systems.
This work addresses the challenges of human activity recognition in smart homes, where sparse sensor signals and similar local patterns hinder accurate modeling of semantically complex daily behaviors. To overcome these limitations, the authors propose TRACE, a framework that reframes activity recognition as a context-aware temporal reasoning task rather than isolated local classification. By integrating multi-source sensor evidence with user-specific contextual priors, TRACE enables coherent and robust semantic inference. The approach effectively mitigates prediction fragmentation and significantly improves recognition accuracy for complex activities on both public benchmarks and real-world deployments. Furthermore, it demonstrates consistent robustness under cross-domain scenarios and in the presence of missing modalities.
This work addresses the limitations of existing reminder systems in leveraging the multimodal sensing capabilities of smart homes and the lack of natural means for users to express complex, context-aware reminders. We propose a novel framework based on natural language and conversational interaction that enables users to flexibly specify multidimensional reminder intents—including temporal, activity-based, sensor-derived, and state-dependent conditions—using everyday language. Through guided dialogue, the system structures ambiguous user expressions into executable rules. The architecture integrates natural language understanding, context-aware computation, sensor fusion, and a rule-based reasoning engine. Two user studies (N=40 and N=10) demonstrate that our approach effectively handles diverse and complex reminder logic, significantly improving alignment between user intent and system interpretation.
This study addresses the challenge of obtaining fine-grained annotations for wearable activity data in everyday environments, where retrospective labeling is costly and temporal boundaries are often ambiguous. The authors propose an in-situ annotation approach based on user-defined trigger-action rules: when an audio event occurs, the system prompts the user to provide open-vocabulary activity labels in real time, while simultaneously capturing multimodal sensor data with precise temporal boundaries. By transforming users into active participants, this method integrates opportunistic crowd-sensing with natural language labeling, substantially reducing annotation burden and enhancing ecological validity. Experimental results demonstrate high system log accuracy (97.30% recall, 97.15% precision), and expert evaluations confirm its superior alignment with real-world scenarios, with a field pilot further validating its feasibility in free-living conditions.
This study addresses the challenge of generating structured event logs from multimodal data such as videos to support business process mining. The authors propose an end-to-end approach that first maps video frames into feature vectors using image embeddings, then performs temporal segmentation via an inter-frame similarity matrix. Subsequently, a generalized few-shot classification method automatically assigns semantic labels to the resulting segments, yielding a timestamped, structured event sequence. This work represents the first integration of image embeddings with few-shot learning for the automatic transformation of raw video into process-mining-ready event logs, thereby overcoming the traditional reliance on pre-structured input data. The method’s effectiveness and practicality are validated through experiments in real-world scenarios.
This work addresses the lack of an event-centered evaluation framework in existing research on proactive agents, where open-ended tasks pose significant assessment challenges. The authors propose the first event-oriented benchmark for evaluating proactive assistance, built upon synthetic yet realistic multi-threaded, noisy, and dynamic instant messaging data. The framework introduces evaluation dimensions including response timing and correctness across single- and multi-step interactions. Leveraging large language models within a task pipeline, the system performs event detection, spatiotemporal information extraction, and contextually appropriate proactive response generation. Experiments across eight mainstream large language models reveal pervasive issues such as over-responsiveness and difficulty handling event cancellations; even the best-performing model, GPT-5.1, achieves correct behavior in only 26.7% of scenarios, underscoring both the benchmark’s rigor and its necessity for advancing proactive agent capabilities.