intent-guided sequence modeling

Design and build sequence-modeling systems that infer, represent, and predict structured user intents from sequential interactions and attribute signals (intent recognition and attribute-based intent modeling). Integrate those structured intent representations or priors—personal and public—into downstream predictors (for example, recommendation models) to guide predictions and improve accuracy and interpretability.

intent-guidedsequencemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the challenge that emerging user intents—arising from evolving interaction sequences—undermine the effectiveness of conventional sequential recommendation models, this paper proposes the Incremental Multi-Intent Adaptive framework (IMA) and its enhanced variant, Elastic Multi-Intent Adaptation (EMA). Methodologically, we introduce a novel dynamic intent modeling mechanism that synergistically integrates capsule networks with self-attention, incorporating intent preservation, emergent-intent detection, and projection-based pruning modules, alongside an intent activity assessment module to jointly optimize intent growth and forgetting. Technically, our approach pioneers the integration of incremental learning, intent vector projection compression, and dynamic activity modeling. Extensive experiments on multiple real-world datasets demonstrate a 12.6% improvement in Recall@20; under memory constraints, historical intent recognition accuracy remains above 98%, significantly outperforming state-of-the-art baselines. This work is the first to systematically resolve the problem of incremental multi-intent sequential recommendation.

Adapting to new user intents in sequential recommendationElastically managing intents under memory constraintsPreventing forgetting existing user intents during updates

Small Models, Big Results: Achieving Superior Intent Extraction through Decomposition

Sep 15, 2025
DC
Danielle Cohen
🏛️ Google | Bar-Ilan University

Resource-constrained edge devices face challenges in accurately understanding user intent from UI interaction traces, while simultaneously ensuring privacy preservation and real-time responsiveness. Method: This paper proposes a two-stage decomposed architecture: (1) generating structured sequential summaries of interaction behaviors, followed by (2) lightweight intent inference based on these summaries. The approach integrates context-aggregated enhancement and task-adaptive fine-tuning to strengthen semantic modeling capabilities of small models. Contribution/Results: Experimental results demonstrate that, under identical privacy guarantees and low-latency constraints, the proposed method achieves higher intent recognition accuracy than state-of-the-art large multimodal language models. It establishes an efficient, privacy-aware, and real-time interaction understanding paradigm for on-device intelligent agents.

Enhancing intent understanding through decomposed structured summarizationImproving intent extraction accuracy in small on-device modelsOvercoming limitations of privacy-preserving low-latency models

This study addresses the challenge of reliably maintaining user goal consistency across multiple models, languages, and prompting frameworks. The authors propose a protocol-like communication layer grounded in structured intent representations—specifically the 5W3H schema—and systematically evaluate its cross-lingual and cross-domain alignment efficacy on Claude, GPT-4o, and Gemini 2.5 Pro. Leveraging both automated evaluation with DeepSeek-V3 and a user study involving 50 participants, the results demonstrate that structured prompting substantially reduces cross-lingual goal drift, lowering the standard deviation of alignment scores from 0.470 to 0.020. The findings further reveal a weak-model compensation effect and the critical role of dimensional decomposition. Notably, Gemini exhibits a +1.006 improvement in goal alignment, accompanied by a 60% reduction in interaction turns and an increase in user satisfaction from 3.16 to 4.04.

cross-model robustnessgoal alignmenthuman-AI interaction

This work addresses key limitations in existing intent-aware recommender systems, which typically rely on a predefined number of intents, are sensitive to behavioral sequence quality, and lack explicit semantic grounding, resulting in coarse-grained intent representations. To overcome these issues, the authors propose a novel approach that leverages sparse autoencoders to unsupervisedly disentangle fine-grained, interpretable intent spaces from text embeddings generated by large language models, eliminating the need to predefine the number of intents. The method distinguishes between user-specific personalized intents and cross-user common intents and introduces a multi-branch attention mechanism that adaptively integrates temporal dynamics with intent priors. Extensive experiments on multiple public datasets demonstrate significant performance gains over state-of-the-art baselines, while also yielding human-interpretable explanations for recommendations.

intent-based recommendationinterpretable intentssemantic grounding

Beyond Item Dissimilarities: Diversifying by Intent in Recommender Systems

May 20, 2024
YW
Yuyan Wang
🏛️ Stanford University | Google | Google DeepMind

Long-term user experience degradation in recommender systems remains a critical challenge. Method: This paper proposes a full-page diversity optimization framework driven by cross-session stable user intents, abandoning conventional item-similarity-based diversity control. Its core innovation is the first integration of real-time dynamic intent prediction into the final ranking stage, employing a Bayesian belief updating mechanism to model and balance multi-intent representations at the intent level—not the item level—combined with serialized position-aware ranking for online intent inference. Contribution/Results: Deployed in YouTube’s live production environment, the method significantly improves DAU and user satisfaction. Empirical evidence demonstrates that intent-driven diversity effectively enhances long-term user retention and experiential enjoyment.

Long-term User SatisfactionReal-time User PreferencesRecommendation Diversity

Latest Papers

What's happening recently
View more

This study investigates whether a structured intent representation based on the 5W3H framework can effectively generalize across multiple languages (Chinese, English, and Japanese) and diverse large language models to enhance intent alignment and accessibility. Leveraging the PPS framework, the authors conduct 2,160 cross-lingual and cross-model controlled experiments under four prompting conditions. They demonstrate for the first time that AI-generated 5W3H prompts perform comparably to human-crafted ones, significantly reducing user input burden. The work also uncovers a “dual inflation bias” in unstructured prompts, whose deceptively low output variance misrepresents model behavior. Results show that structured prompting not only improves target alignment but also reshapes and more accurately reflects cross-model output variance, offering a novel paradigm for prompt engineering.

5W3H promptingcross-language generalizationcross-model consistency

Intent-Guided Reasoning for Sequential Recommendation

Dec 15, 2025
YS
Yifan Shao
🏛️ The Chinese University of Hong Kong | Hong Kong University of Science and Technology (Guangzhou)

Existing sequential recommendation models suffer from two key limitations: inference instability (i.e., sensitivity to behavioral noise) and superficiality (i.e., modeling only item-level transitions while neglecting deeper behavioral patterns). To address these, we propose an intent-anchored two-stage reasoning framework. First, a latent intent distiller explicitly captures high-order user intents; second, an intent-aware deliberative reasoner enables stable, deep-level inference. Our method incorporates a dual-attention decoupling architecture, a frozen encoder augmented with learnable intent tokens, multi-view intent consistency regularization, and a noise-robust training strategy. Extensive experiments show an average 7.13% improvement over baselines across three public benchmarks; under 20% behavioral noise, performance degrades by only 10.4%, substantially outperforming state-of-the-art methods. Our core contribution is the first explicit use of high-order user intent as a stable reasoning anchor—uniquely balancing robustness, interpretability, and deep behavioral pattern modeling.

Existing methods exhibit surface-level reasoning by memorizing item transitions.Models lack understanding of intrinsic user behavior patterns and intents.Sequential recommendation systems suffer from reasoning instability due to noise.

This work addresses the limitations of traditional sequential recommendation models, which represent user behavior as a single sequence and consequently suffer from mixed multi-interest signals and contextual contamination, hindering the capture of high-intent actions. To overcome this, the authors propose a constructive multi-sequence learning framework that employs a learnable sequence construction module to explicitly disentangle user history in latent space, generating thematically coherent subsequences. Coupled with a linear attention mechanism, this approach enables efficient and focused modeling. Departing from the conventional single-sequence paradigm, the method introduces the concept of “context engineering” to achieve clean, disentangled representations of multiple user interests. The framework has been deployed across four core recommendation scenarios at Meta, demonstrating significant improvements in recommendation accuracy and model focus for both ranking and retrieval tasks.

context pollutionmulti-faceted interestsrecommendation systems

Language Models as Semantic Augmenters for Sequential Recommenders

Oct 20, 2025
MV
Mahsa Valizadeh
🏛️ Texas A&M University

To address the limited representational capacity of sequential recommendation models caused by sparse semantic context, this paper proposes LaMAR—a data-centric framework that pioneers the use of large language models (LLMs) as semantic enhancers. LaMAR automatically generates multi-dimensional semantic signals (e.g., usage scenarios, item intents, and topical summaries) for user behavior sequences under few-shot settings. It integrates item metadata into controllable prompt engineering to ensure high novelty and diversity in signal generation, and seamlessly embeds these signals into mainstream sequential recommendation architectures. Extensive experiments on multiple benchmark datasets demonstrate that LaMAR significantly improves recommendation performance—achieving an average 12.7% gain in Recall@20. Ablation studies confirm that the generated semantic signals effectively strengthen downstream models’ semantic understanding and generalization capability.

Enriching sequential interaction data with semantic contextGenerating auxiliary signals from user intent and item relationshipsImproving performance of sequential recommenders through LLM augmentation

This work addresses key challenges in modeling human daily activity sequences—namely, long-tailed distributions, limited interpretability, and the difficulty of unifying diverse tasks—by proposing a Behavior Understanding Alignment (BUA) framework. BUA is the first approach to integrate large language models (LLMs) into behavior modeling through structured curriculum learning, leveraging sequence representations from pretrained behavior models as alignment anchors. By combining a three-stage curriculum with a multi-turn dialogue mechanism, the framework unifies behavior prediction and generation within a single architecture. BUA effectively bridges the semantic gap between behavioral data and natural language modalities, significantly outperforming existing methods on two real-world datasets while offering strong interpretability and robust multi-task generalization capabilities.

behavior predictionhuman behavior modelinglarge language models

Hot Scholars

DS

Daniel S. Brown

Assistant Professor, Robotics Center and Kahlert School of Computing, University of Utah
🏆 Reward Learning🛡 Safe and Robust AI🤖 Robot Learning✋ Human-Robot Interaction
HY

Hanchen Yang

Georgia Institute of Technology
Computer ArchitectureMachine Learning
MA

Mohammad Aliannejadi

Assistant Professor of Computer Science. IRLab, University of Amsterdam
Information RetrievalNatural Language ProcessingMachine Learning
QZ

Qing Zong

HKUST
Natural Language ProcessingLarge Language ModelsFactualityUncertainty Calibration
CL

Chenyi Lei

Kuaishou Technology
Recommender SystemInformation RetrievalGenerative RecommendationMultimodal