Score
Design and build representations and interpreters that map heterogeneous natural-language statements of goals, desires, and preferences into a structured intent space (typically numeric vectors or latent dimensions). This includes methods to extract fine‑grained intents from text, disentangle intent semantics from noise, interpret latent dimensions as intents, normalize and reconcile personal and public intents across contexts, and quantify or rank preference priorities for downstream retrieval, scoring, or decision modules.
This paper addresses two core challenges in natural language understanding: commonsense reasoning and intent detection. Through a cross-disciplinary literature review of 28 representative papers published at ACL, EMNLP, and CHI between 2020–2025, we apply systematic classification and thematic analysis—uniquely integrating NLP and HCI perspectives for the first time. We identify four commonsense reasoning pathways (zero-shot learning, cultural adaptation, structured evaluation, and interactive modeling) and four intent detection paradigms (open-set modeling, generative identification, unsupervised clustering, and human-centered design). Our analysis reveals three critical gaps: insufficient semantic grounding, weak cross-scenario generalization, and inconsistent benchmarking practices. To bridge these, we propose an evolutionary framework centered on adaptability, multilinguality, and context awareness. This framework provides both theoretical foundations and actionable guidelines for developing next-generation AI systems that are interpretable, robust, and human-centered.
Traditional unimodal text-based intent recognition suffers from limited contextual expressiveness, while human–computer interaction increasingly demands robust integration of heterogeneous signals. This paper systematically surveys deep learning–based multimodal intent recognition, focusing on synergistic modeling of textual, audio, visual, and physiological modalities. It traces the technical evolution from unimodal baselines to cross-modal fusion, emphasizing breakthrough applications of Transformer architectures in cross-modal alignment, feature fusion, and representation learning. We catalog 12 mainstream multimodal datasets, unify evaluation metrics, and identify representative application scenarios. A three-dimensional taxonomy—spanning modality combinations, fusion levels (early/late/hybrid), and learning paradigms (supervised/self-supervised/few-shot)—is proposed. Key challenges—including modality asynchrony, few-shot generalization, and model interpretability—are critically analyzed. Future directions include optimized cross-modal alignment, neuro-symbolic integration, and edge-efficient lightweight modeling, offering a structured reference for advancing multimodal intent understanding.
To address the challenges of fine-grained multimodal intent modeling, insufficient cross-modal alignment, and severe noise interference in recommender systems, this paper proposes a large language model (LLM)-based multimodal intent representation framework. Methodologically, we design a dual-tower neural architecture that jointly encodes user behavioral sequences and heterogeneous textual signals (e.g., reviews and item descriptions). We introduce a novel synergistic alignment mechanism combining pairwise contrastive learning and translation-style cross-modal mapping, augmented by momentum-based knowledge distillation to mitigate modality-specific representation heterogeneity and noise. Extensive experiments on three public benchmarks demonstrate significant improvements in recommendation accuracy—achieving an average +3.2% gain in NDCG@10—and enhanced interpretability of learned user intents. The implementation code is publicly available.
This work addresses key limitations in existing intent-aware recommender systems, which typically rely on a predefined number of intents, are sensitive to behavioral sequence quality, and lack explicit semantic grounding, resulting in coarse-grained intent representations. To overcome these issues, the authors propose a novel approach that leverages sparse autoencoders to unsupervisedly disentangle fine-grained, interpretable intent spaces from text embeddings generated by large language models, eliminating the need to predefine the number of intents. The method distinguishes between user-specific personalized intents and cross-user common intents and introduces a multi-branch attention mechanism that adaptively integrates temporal dynamics with intent priors. Extensive experiments on multiple public datasets demonstrate significant performance gains over state-of-the-art baselines, while also yielding human-interpretable explanations for recommendations.
Resource-constrained edge devices face challenges in accurately understanding user intent from UI interaction traces, while simultaneously ensuring privacy preservation and real-time responsiveness. Method: This paper proposes a two-stage decomposed architecture: (1) generating structured sequential summaries of interaction behaviors, followed by (2) lightweight intent inference based on these summaries. The approach integrates context-aggregated enhancement and task-adaptive fine-tuning to strengthen semantic modeling capabilities of small models. Contribution/Results: Experimental results demonstrate that, under identical privacy guarantees and low-latency constraints, the proposed method achieves higher intent recognition accuracy than state-of-the-art large multimodal language models. It establishes an efficient, privacy-aware, and real-time interaction understanding paradigm for on-device intelligent agents.
This paper addresses the insufficient modeling of textual implicit semantics by proposing a novel paradigm that explicitly transforms “subtext” into a verifiable set of propositions. Methodologically, it systematically leverages large language models to generate implicit inference propositions from text, which are then validated for plausibility by human annotators; the resulting proposition semantics are subsequently fused into the original text representations. Key contributions include: (1) introducing the first annotated framework for implicit propositions tailored to social science tasks; and (2) demonstrating that this explicit modeling significantly outperforms literal-only representations across three distinct tasks—argument similarity assessment, public opinion interpretation, and legislative behavior simulation—with average improvements of 12.7% in F1 or accuracy. Results substantiate the effectiveness and generalizability of structured implicit semantic modeling for enhancing human-like semantic understanding.
To address two key limitations of large language models (LLMs) in multi-intent spoken language understanding (SLU)—word-level alignment bias arising from autoregressive generation and insufficient modeling of semantic-level intent relationships—this paper proposes the EN-LLM and ENSI-LLM architectures. Methodologically, it introduces: (1) a restructured entity-slot modeling paradigm; (2) the Sub-Intent Instruction (SII) mechanism, the first of its kind to explicitly guide fine-grained intent decomposition; (3) synthetic benchmarks LM-MixATIS and LM-MixSNIPS; and (4) two novel fine-grained evaluation metrics—Entity Slot Accuracy (ESA) and Combined Semantic Accuracy (CSA). Experiments demonstrate state-of-the-art performance on multi-intent SLU tasks, with significant improvements in robustness to complex intent compositions and distributional shifts, as well as enhanced generalization capability.
This study addresses the challenge of reliably maintaining user goal consistency across multiple models, languages, and prompting frameworks. The authors propose a protocol-like communication layer grounded in structured intent representations—specifically the 5W3H schema—and systematically evaluate its cross-lingual and cross-domain alignment efficacy on Claude, GPT-4o, and Gemini 2.5 Pro. Leveraging both automated evaluation with DeepSeek-V3 and a user study involving 50 participants, the results demonstrate that structured prompting substantially reduces cross-lingual goal drift, lowering the standard deviation of alignment scores from 0.470 to 0.020. The findings further reveal a weak-model compensation effect and the critical role of dimensional decomposition. Notably, Gemini exhibits a +1.006 improvement in goal alignment, accompanied by a 60% reduction in interaction turns and an increase in user satisfaction from 3.16 to 4.04.
Current holistic evaluation approaches struggle to distinguish between structural replication and fidelity to user intent in large language model outputs. This work proposes a dimension-level intent fidelity assessment framework that, through structured prompt ablation, human evaluation, and weight perturbation, reveals for the first time a systematic divergence between structural fidelity and intent fidelity. Experiments show that among high-scoring outputs in both Chinese and English, 25.7% and 58.6%, respectively, exhibit deficiencies in dimensional intent alignment. Moreover, dimension-level scores demonstrate significantly higher agreement with human judgments than holistic scores, offering a more precise reflection of output quality flaws.
This work addresses the limitations of explicit symbolic spaces in language models—namely linguistic redundancy, discretization bottlenecks, sequential inefficiency, and semantic loss—which constrain computational efficiency and expressive capacity. The study systematically reviews advances in latent space research and introduces, for the first time, a five-dimensional analytical framework tailored to language models: foundations, evolution, mechanisms, capabilities, and outlook. This unified framework integrates architecture, representation, computation, and optimization, while linking technical pathways to higher-order abilities such as reasoning, memory, and embodiment. By synthesizing and categorizing cutting-edge research, the paper elucidates the pivotal role of latent spaces in enhancing model efficiency and capability, and clearly identifies key challenges and promising directions for future work.
This work addresses the semantic gap between short user queries and product semantic identifiers (SIDs) in e-commerce search, along with stringent latency constraints. To this end, the authors propose CaLIR, a framework that performs coarse-to-fine implicit intent reasoning guided by category hierarchies, learning continuous latent states to align user shopping intent with SIDs without relying on explicit chain-of-thought reasoning that incurs high latency. CaLIR innovatively integrates query-aware intent path modeling, dynamic prefix Trie-constrained decoding, and hierarchical semantic reasoning, achieving substantial retrieval gains without compromising efficiency. Experiments demonstrate that CaLIR strikes a superior balance between retrieval effectiveness and inference efficiency across multilingual e-commerce datasets, while exhibiting strong transferability and robustness across diverse category taxonomies and generative backbone models.
This work addresses the challenges of discovering user intents in unlabeled multimodal dialogues, where the absence of explicit supervision and limited model interpretability hinder performance. The authors propose MCSP, a novel approach that uniquely integrates multimodal large language model (MLLM)-guided contrastive reasoning with semantic propagation. Specifically, high-level semantic concepts generated by an MLLM guide unsupervised clustering, and these semantics are further refined through propagation over a semantically weighted graph to enforce both local consistency and global semantic alignment. Evaluated on three multimodal intent datasets, MCSP significantly outperforms state-of-the-art methods while yielding semantically interpretable clustering outcomes.