SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Evaluating smartphone personal assistants requires privacy-sensitive data with rich contextual grounding and ground-truth answers, yet real user traces are rarely shareable. To address this, this work proposes a digital twin simulation system grounded in physical laws and event provenance, which integrates real-world maps, weather, holidays, and network data to drive virtual users through a full simulated day, automatically generating evaluation data with constructively defined ground truth. Innovatively, the system uses snapshot pointers as labels, eliminating the need for post-hoc annotation or large-model-based judgments, thereby ensuring privacy preservation, reproducibility, and label determinism. The synthesized data closely matches real-user behavior in both category distribution (JSD 0.070) and communication timing (JSD < 0.1), and successfully uncovers 78 retrieval-related failures in a commercial assistant, predominantly involving call and SMS records.
📝 Abstract
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.
Problem

Research questions and friction points this paper is trying to address.

smartphone personal assistants
evaluation data
privacy-sensitive
context-rich
ground truth
Innovation

Methods, ideas, or system contributions that make the work stand out.

digital twin
event-sourced simulation
ground-truth evaluation data
privacy-preserving benchmarking
deterministic persona modeling
🔎 Similar Papers
No similar papers found.