multimodal user profiling

Designs and builds representations, models, and data pipelines that aggregate signals from multiple modalities into per-user profiles and timelines, inferring stable and transient preferences, ethical stances, and other attributes. It also analyzes longitudinal and sequential behavior to capture temporal dependencies and trajectories, segments users by behavior or preference, and produces profile summaries and embeddings for downstream analysis or systems.

multimodaluserprofiling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Exploring temporal dynamics in digital trace data: mining user-sequences for communication research

May 24, 2025
YF
Yangliu Fan
🏛️ Weizenbaum Institute for Networked Society | Free University of Berlin

Communication research has long faced a methodological tension between static analytical approaches and the inherently dynamic nature of diffusion processes, hindering fine-grained temporal modeling of digital traces. To address this, we propose “hyper-longitudinal analysis”—the first systematic, user-level dynamic diffusion framework centered on raw, unaggregated time series, preserving full temporal structure without dimensional reduction. Integrating six computational paradigms—sequence analysis, process mining, language modeling, temporal pattern discovery, trajectory clustering, and behavioral modeling—we establish a theory-guided, computationally synergistic pipeline. Validated on 1.26 million donation-related digital traces from 309 users, our approach uncovers three key empirical regularities: (1) recurrent periodicity in sharing behavior, (2) path-dependent propagation mechanisms, and (3) heterogeneous individual-level temporal patterns. The framework provides a scalable, generalizable methodological foundation for dynamic communication modeling.

Bridging the gap between theoretical dynamics and non-dynamical methodologies in communication researchDeveloping a computational framework to analyze time-evolving user-sequences from digital trace dataEnhancing understanding of temporal dimensions in communication processes using high-resolution data

How Can Time Series Analysis Benefit From Multiple Modalities? A Survey and Outlook

Mar 14, 2025
HL
Haoxin Liu
🏛️ Georgia Institute of Technology | Bytedance Inc. | Squirrel AI | The University of Virginia | Cornell University

Multimodal fusion in time-series analysis (TSA) remains underexplored, with no systematic survey or established paradigm. Method: We introduce “Multimodal-Empowered Time-Series Analysis” (MM4TSA), proposing a three-tier benefit framework—foundation model reuse, multimodal extended modeling, and cross-modal interactive learning—and categorize existing methods by modality (e.g., text, image, audio). Leveraging techniques including transfer learning, cross-modal alignment, and foundation model adaptation, we conduct a systematic literature review to identify three critical gaps: modality selection, heterogeneous modality combination, and task generalization. Contribution/Results: We release the first dynamic MM4TSA GitHub knowledge base, clarifying technical lineages and offering an extensible methodology guide. This work establishes a foundational taxonomy and fosters the evolution of TSA toward cross-modal collaborative paradigms.

Explores benefits of multiple modalities in time series analysis.Identifies gaps and future opportunities in MM4TSA research.Reviews reusing, extending, and interacting modalities for TSA.

This work addresses the limitations of existing user simulation methods, which often treat users as static entities or rely on overly generalized historical contexts, thereby failing to capture the dynamic evolution of individual behavior. To overcome this, the paper proposes TWICE, a novel framework that introduces a life-event-driven memory mechanism to model how users’ past experiences shape their current expressions. TWICE integrates structured user profiles with a two-stage generation process—decoupling content planning from stylistic adaptation—to enable fine-grained modeling of behavioral dynamics. Built upon large language models, TWICE demonstrates significant improvements over strong baselines on a large-scale longitudinal Twitter dataset, achieving superior performance across comprehensive metrics including authenticity, consistency, and human-likeness.

behavioral dynamicsevent-driven modelingpersonalized behavior

Existing behavioral modeling approaches treat user actions as discrete event sequences, neglecting the contextual information embedded in inter-action time intervals—leading to incomplete behavioral understanding and poor interpretability. To address this, we propose the dual-scale Action-Timing Context (ATC) framework, the first to systematically model *inter-action temporal context* by jointly embedding action types and time intervals within a unified representation space, thereby capturing fine-grained temporal structure. ATC employs dual-scale temporal embedding and low-dimensional action representation learning to yield interpretable and consistent behavioral embeddings. Extensive experiments on multiple real-world digital platform log datasets demonstrate that ATC significantly improves performance in behavioral prediction, post-hoc interpretability, and analysis of sociological mechanisms—including knowledge accumulation and information diffusion—thereby filling a critical gap in temporal structural modeling of human behavior.

Embedding actions and time intervals jointlyModeling inter-temporal context in human actionsUnderstanding human behavior on digital platforms

PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits

Sep 14, 2025
LL
Loka Li
🏛️ Mohamed bin Zayed University of Artificial Intelligence | University of California San Diego | Carnegie Mellon University | Australian National University

Existing research lacks large-scale multimodal personality datasets integrating behavioral descriptors, facial images, and biographical information, hindering cross-modal modeling of human behavioral traits. Method: We introduce PersonaX—a scalable multimodal dataset comprising CelebPersona and AthlePersona—covering over 10,000 public figures and athletes with behavioral trait annotations, facial images, and structured biographical data. We propose Causal Representation Learning (CRL), a theoretically identifiable causal inference framework for multimodal and multi-measurement settings. CRL jointly processes textual, visual, and structured data using three state-of-the-art large language models and validates causal relationships via statistical independence tests. Contribution/Results: Empirical evaluation on synthetic and real-world data confirms robust cross-modal associations between facial/biographical features and behavioral traits. PersonaX establishes the first reproducible, extensible, causally grounded multimodal benchmark for personalized AI and computational social science.

Analyzing relationships between LLM-inferred traits and biographical featuresDeveloping causal representation learning for multimodal trait analysisMultimodal datasets combining behavioral traits with facial attributes

Latest Papers

What's happening recently
View more

Traditional user profiling approaches rely on discriminative models and manual feature engineering, struggling to capture long-tail behaviors and often yielding fragmented, logically inconsistent profiles. This work proposes UserGPT, a novel framework that leverages large language models to transform massive, noisy user behavioral logs into coherent user narratives, enabling holistic personality inference. UserGPT introduces a dual-path paradigm—combining attribute generation and summary generation—and integrates a user behavior simulation engine, a data semanticization module, multi-stage supervised fine-tuning, and a dual-filter grouped relative policy optimization (DF-GRPO) strategy. Evaluated on HPR-Bench, the framework achieves an Avg@10 of 0.7325 for label prediction and an Acc_Ex of 0.7528 for summary generation, while compressing behavioral records by 97.9% without significant loss of critical information.

behavioral traceslong-tail behaviorspersona reasoning

This work addresses the limitations of existing personalized language models, which predominantly rely on explicit user preferences and struggle to infer authentic interests from naturally occurring multimodal social media traces. To bridge this gap, we introduce SocialPersona—the first benchmark for personalized evaluation grounded in real longitudinal social media data encompassing text, images, and timestamps. SocialPersona enables structured user profiling and personalized dialogue generation by distinguishing between stable and recent interests, while revealing the complementary roles of textual and visual signals in preference inference. Experiments demonstrate that current multimodal large language models can recognize broad interests but exhibit limited capability in modeling fine-grained and recent preferences; performance further degrades when leveraging these profiles for dialogue generation, underscoring long-term cross-modal user modeling as a critical open challenge.

long-horizon inferencemultimodal social-media contextpersonalized profiling

Current AI systems predominantly focus on representing user behaviors while overlooking the underlying motivations, thereby failing to address users’ deeper needs. This work proposes a novel “behavior latticing” architecture that, for the first time, models semantic relationships among discrete interaction behaviors using a lattice structure. By integrating multi-turn unstructured interaction data, the approach enables cross-behavioral inference of latent user motivations, uncovering implicit needs even users themselves may not recognize, and generating interpretable insights. Experimental results demonstrate that the proposed method significantly outperforms existing techniques in both motivation identification accuracy and depth of interpretability. The framework has been successfully deployed in personal AI agents, enhancing immediate utility while better aligning with long-term user value.

behavior understandingpersonal AIunstructured interactions

Existing social simulation approaches predominantly rely on demographic attributes, which inadequately capture individual psychological consistency and thereby limit the fidelity of intervention-effect modeling. This work proposes SPIRIT, a novel framework that introduces semi-structured personality representations—integrating structured psychological traits with unstructured narrative text—into large-scale agent-based simulations to drive context-sensitive, individualized opinion and behavioral responses via large language models. Evaluated on Ipsos KnowledgePanel data, SPIRIT significantly outperforms demographic-only baselines, not only reproducing self-reported survey responses with higher accuracy but also effectively capturing the heterogeneity inherent in human reactions. This advancement marks a paradigm shift from static prediction to dynamic, psychologically grounded simulation of social behavior.

human opinion modelingindividualized trajectoriespersona-based simulation

Hot Scholars

KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
XH

Xiangnan He

University of Science and Technology of China
RecommendationCausalityBig DataInformation Retrieval
GZ

Guorui Zhou

Unknown affiliation
Recommender System,Advertising,Artificial Intelligence,Machine Learning,NLP
MS

Milad Sabouri

DePaul University
Machine LearningReinforcement LearningRecommender Systems
EF

Emilio Ferrara

Professor of Computer Science at the University of Southern California
Human-Centered AISocial ComputingNetwork ScienceAI Safety