Semantic-aware Graph-guided Behavior Sequences Generation with Large Language Models for Smart Homes

📅 2025-08-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing models for smart home behavior analysis, trained on static datasets, exhibit poor generalization under behavioral drift caused by seasonal changes and evolving user habits; re-collecting and annotating new data is costly, time-consuming, and raises privacy concerns. Method: We propose SmartGen—a novel framework integrating time- and semantic-aware subsequence partitioning, latent-space behavioral clustering compression, graph-guided sequence generation, and a two-stage anomaly filtering mechanism—to leverage large language models (LLMs) for synthesizing high-fidelity, context-aware user behavior sequences that enable continual downstream model adaptation. Contribution/Results: Evaluated on three real-world smart home datasets, SmartGen achieves average improvements of 85.43% in anomaly detection and 70.51% in behavior prediction over state-of-the-art methods, demonstrating superior robustness to behavioral drift without requiring new labeled data.

Technology Category

Humans and AI: Human-Aware Planning and Behavior PredictionNatural Language Processing: GenerationIntelligent Robots: Behavior Learning & Control

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
As smart homes become increasingly prevalent, intelligent models are widely used for tasks such as anomaly detection and behavior prediction. These models are typically trained on static datasets, making them brittle to behavioral drift caused by seasonal changes, lifestyle shifts, or evolving routines. However, collecting new behavior data for retraining is often impractical due to its slow pace, high cost, and privacy concerns. In this paper, we propose SmartGen, an LLM-based framework that synthesizes context-aware user behavior data to support continual adaptation of downstream smart home models. SmartGen consists of four key components. First, we design a Time and Semantic-aware Split module to divide long behavior sequences into manageable, semantically coherent subsequences under dual time-span constraints. Second, we propose Semantic-aware Sequence Compression to reduce input length while preserving representative semantics by clustering behavior mapping in latent space. Third, we introduce Graph-guided Sequence Synthesis, which constructs a behavior relationship graph and encodes frequent transitions into prompts, guiding the LLM to generate data aligned with contextual changes while retaining core behavior patterns. Finally, we design a Two-stage Outlier Filter to identify and remove implausible or semantically inconsistent outputs, aiming to improve the factual coherence and behavioral validity of the generated sequences. Experiments on three real-world datasets demonstrate that SmartGen significantly enhances model performance on anomaly detection and behavior prediction tasks under behavioral drift, with anomaly detection improving by 85.43% and behavior prediction by 70.51% on average. The code is available at https://github.com/horizonsinzqs/SmartGen.
Problem

Research questions and friction points this paper is trying to address.

Generates synthetic user behavior data for smart homes
Addresses behavioral drift in static smart home models
Enhances anomaly detection and behavior prediction accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time and Semantic-aware Split module
Semantic-aware Sequence Compression
Graph-guided Sequence Synthesis
Z
Zhiyao Xu
Tsinghua Shenzhen International Graduate School, Shenzhen, China
D
Dan Zhao
Peng Cheng Laboratory, Shenzhen, China
Q
Qingsong Zou
Tsinghua Shenzhen International Graduate School, Peng Cheng Laboratory, Shenzhen, China
Q
Qing Li
Peng Cheng Laboratory, Shenzhen, China
Y
Yong Jiang
Tsinghua Shenzhen International Graduate School, Peng Cheng Laboratory, Shenzhen, China
Y
Yuhang Wang
Southwest University, Chongqing, China
Jingyu Xiao
Jingyu Xiao
Tsinghua University
Data MiningLarge Language ModelsComputer NetworkMLLM4Code