Score
Designs and builds representations, models, and data pipelines that aggregate signals from multiple modalities into per-user profiles and timelines, inferring stable and transient preferences, ethical stances, and other attributes. It also analyzes longitudinal and sequential behavior to capture temporal dependencies and trajectories, segments users by behavior or preference, and produces profile summaries and embeddings for downstream analysis or systems.
Traditional recommender systems are constrained by single-task, single-scenario, single-modality, and single-behavior modeling, limiting their ability to capture users’ dynamic and complex preferences. To address this, we propose the first unified four-dimensional joint modeling paradigm—integrating multi-task learning, multi-scenario adaptation, multimodal fusion, and multi-behavior modeling. We systematically survey key techniques, including deep neural architectures, transfer learning, multi-source feature integration, behavioral sequence modeling, and cross-domain representation alignment, distilling common architectural principles and training strategies. Further, we construct a structured, knowledge-graph-inspired taxonomy that clarifies capability boundaries and applicability conditions of existing methods. For the first time, we deliver a reusable methodology guide and an open-problem checklist, bridging theoretical foundations with practical implementation. This work provides both principled guidance for algorithm design and actionable pathways for industrial deployment.
Communication research has long faced a methodological tension between static analytical approaches and the inherently dynamic nature of diffusion processes, hindering fine-grained temporal modeling of digital traces. To address this, we propose “hyper-longitudinal analysis”—the first systematic, user-level dynamic diffusion framework centered on raw, unaggregated time series, preserving full temporal structure without dimensional reduction. Integrating six computational paradigms—sequence analysis, process mining, language modeling, temporal pattern discovery, trajectory clustering, and behavioral modeling—we establish a theory-guided, computationally synergistic pipeline. Validated on 1.26 million donation-related digital traces from 309 users, our approach uncovers three key empirical regularities: (1) recurrent periodicity in sharing behavior, (2) path-dependent propagation mechanisms, and (3) heterogeneous individual-level temporal patterns. The framework provides a scalable, generalizable methodological foundation for dynamic communication modeling.
Multimodal fusion in time-series analysis (TSA) remains underexplored, with no systematic survey or established paradigm. Method: We introduce “Multimodal-Empowered Time-Series Analysis” (MM4TSA), proposing a three-tier benefit framework—foundation model reuse, multimodal extended modeling, and cross-modal interactive learning—and categorize existing methods by modality (e.g., text, image, audio). Leveraging techniques including transfer learning, cross-modal alignment, and foundation model adaptation, we conduct a systematic literature review to identify three critical gaps: modality selection, heterogeneous modality combination, and task generalization. Contribution/Results: We release the first dynamic MM4TSA GitHub knowledge base, clarifying technical lineages and offering an extensible methodology guide. This work establishes a foundational taxonomy and fosters the evolution of TSA toward cross-modal collaborative paradigms.
This work addresses the limitations of existing user simulation methods, which often treat users as static entities or rely on overly generalized historical contexts, thereby failing to capture the dynamic evolution of individual behavior. To overcome this, the paper proposes TWICE, a novel framework that introduces a life-event-driven memory mechanism to model how users’ past experiences shape their current expressions. TWICE integrates structured user profiles with a two-stage generation process—decoupling content planning from stylistic adaptation—to enable fine-grained modeling of behavioral dynamics. Built upon large language models, TWICE demonstrates significant improvements over strong baselines on a large-scale longitudinal Twitter dataset, achieving superior performance across comprehensive metrics including authenticity, consistency, and human-likeness.
Existing behavioral modeling approaches treat user actions as discrete event sequences, neglecting the contextual information embedded in inter-action time intervals—leading to incomplete behavioral understanding and poor interpretability. To address this, we propose the dual-scale Action-Timing Context (ATC) framework, the first to systematically model *inter-action temporal context* by jointly embedding action types and time intervals within a unified representation space, thereby capturing fine-grained temporal structure. ATC employs dual-scale temporal embedding and low-dimensional action representation learning to yield interpretable and consistent behavioral embeddings. Extensive experiments on multiple real-world digital platform log datasets demonstrate that ATC significantly improves performance in behavioral prediction, post-hoc interpretability, and analysis of sociological mechanisms—including knowledge accumulation and information diffusion—thereby filling a critical gap in temporal structural modeling of human behavior.
Existing research lacks large-scale multimodal personality datasets integrating behavioral descriptors, facial images, and biographical information, hindering cross-modal modeling of human behavioral traits. Method: We introduce PersonaX—a scalable multimodal dataset comprising CelebPersona and AthlePersona—covering over 10,000 public figures and athletes with behavioral trait annotations, facial images, and structured biographical data. We propose Causal Representation Learning (CRL), a theoretically identifiable causal inference framework for multimodal and multi-measurement settings. CRL jointly processes textual, visual, and structured data using three state-of-the-art large language models and validates causal relationships via statistical independence tests. Contribution/Results: Empirical evaluation on synthetic and real-world data confirms robust cross-modal associations between facial/biographical features and behavioral traits. PersonaX establishes the first reproducible, extensible, causally grounded multimodal benchmark for personalized AI and computational social science.
Traditional user profiling approaches rely on discriminative models and manual feature engineering, struggling to capture long-tail behaviors and often yielding fragmented, logically inconsistent profiles. This work proposes UserGPT, a novel framework that leverages large language models to transform massive, noisy user behavioral logs into coherent user narratives, enabling holistic personality inference. UserGPT introduces a dual-path paradigm—combining attribute generation and summary generation—and integrates a user behavior simulation engine, a data semanticization module, multi-stage supervised fine-tuning, and a dual-filter grouped relative policy optimization (DF-GRPO) strategy. Evaluated on HPR-Bench, the framework achieves an Avg@10 of 0.7325 for label prediction and an Acc_Ex of 0.7528 for summary generation, while compressing behavioral records by 97.9% without significant loss of critical information.
This work addresses the limitations of existing personalized language models, which predominantly rely on explicit user preferences and struggle to infer authentic interests from naturally occurring multimodal social media traces. To bridge this gap, we introduce SocialPersona—the first benchmark for personalized evaluation grounded in real longitudinal social media data encompassing text, images, and timestamps. SocialPersona enables structured user profiling and personalized dialogue generation by distinguishing between stable and recent interests, while revealing the complementary roles of textual and visual signals in preference inference. Experiments demonstrate that current multimodal large language models can recognize broad interests but exhibit limited capability in modeling fine-grained and recent preferences; performance further degrades when leveraging these profiles for dialogue generation, underscoring long-term cross-modal user modeling as a critical open challenge.
Current AI systems predominantly focus on representing user behaviors while overlooking the underlying motivations, thereby failing to address users’ deeper needs. This work proposes a novel “behavior latticing” architecture that, for the first time, models semantic relationships among discrete interaction behaviors using a lattice structure. By integrating multi-turn unstructured interaction data, the approach enables cross-behavioral inference of latent user motivations, uncovering implicit needs even users themselves may not recognize, and generating interpretable insights. Experimental results demonstrate that the proposed method significantly outperforms existing techniques in both motivation identification accuracy and depth of interpretability. The framework has been successfully deployed in personal AI agents, enhancing immediate utility while better aligning with long-term user value.
Existing social simulation approaches predominantly rely on demographic attributes, which inadequately capture individual psychological consistency and thereby limit the fidelity of intervention-effect modeling. This work proposes SPIRIT, a novel framework that introduces semi-structured personality representations—integrating structured psychological traits with unstructured narrative text—into large-scale agent-based simulations to drive context-sensitive, individualized opinion and behavioral responses via large language models. Evaluated on Ipsos KnowledgePanel data, SPIRIT significantly outperforms demographic-only baselines, not only reproducing self-reported survey responses with higher accuracy but also effectively capturing the heterogeneity inherent in human reactions. This advancement marks a paradigm shift from static prediction to dynamic, psychologically grounded simulation of social behavior.
研究对比了基于LLM和聚合方法的用户画像策略在流媒体推荐中的效果,探讨了不同策略在准确性、推荐质量及时间窗口设置上的差异。