Score
Design, train, and evaluate models and end-to-end pipelines that map user inputs (queries, utterances, or search requests) to discrete intent labels or structured intent representations; this includes intent taxonomy design, feature and contextual sequence modeling, multi-label detection, and intent extraction. Build and analyze intent-detection and intent-recognition systems — including contextual and query-aware variants — and measure their performance, ambiguity, and error modes.
Traditional unimodal text-based intent recognition suffers from limited contextual expressiveness, while human–computer interaction increasingly demands robust integration of heterogeneous signals. This paper systematically surveys deep learning–based multimodal intent recognition, focusing on synergistic modeling of textual, audio, visual, and physiological modalities. It traces the technical evolution from unimodal baselines to cross-modal fusion, emphasizing breakthrough applications of Transformer architectures in cross-modal alignment, feature fusion, and representation learning. We catalog 12 mainstream multimodal datasets, unify evaluation metrics, and identify representative application scenarios. A three-dimensional taxonomy—spanning modality combinations, fusion levels (early/late/hybrid), and learning paradigms (supervised/self-supervised/few-shot)—is proposed. Key challenges—including modality asynchrony, few-shot generalization, and model interpretability—are critically analyzed. Future directions include optimized cross-modal alignment, neuro-symbolic integration, and edge-efficient lightweight modeling, offering a structured reference for advancing multimodal intent understanding.
This paper addresses two core challenges in natural language understanding: commonsense reasoning and intent detection. Through a cross-disciplinary literature review of 28 representative papers published at ACL, EMNLP, and CHI between 2020–2025, we apply systematic classification and thematic analysis—uniquely integrating NLP and HCI perspectives for the first time. We identify four commonsense reasoning pathways (zero-shot learning, cultural adaptation, structured evaluation, and interactive modeling) and four intent detection paradigms (open-set modeling, generative identification, unsupervised clustering, and human-centered design). Our analysis reveals three critical gaps: insufficient semantic grounding, weak cross-scenario generalization, and inconsistent benchmarking practices. To bridge these, we propose an evolutionary framework centered on adaptability, multilinguality, and context awareness. This framework provides both theoretical foundations and actionable guidelines for developing next-generation AI systems that are interpretable, robust, and human-centered.
To address the challenges of ambiguous intent definitions, high annotation costs, and data scarcity in low-resource language intent classification, this paper proposes a zero-shot retrieval-based intent identification method. Instead of relying on explicit intent labels, the approach models each intent as a set of historically annotated queries and performs intent inference via dense retrieval—matching an input query to its most semantically similar labeled queries in a shared latent space. The method employs a multilingual semantic encoder to produce cross-lingually aligned query embeddings, enhances semantic consistency through contrastive learning, and accelerates retrieval using approximate nearest neighbor (ANN) search. Evaluated on eight low-resource languages, it achieves an average F1 score of 72.4%, outperforming zero-shot fine-tuning baselines by 18.6 percentage points; remarkably, it attains near fully supervised performance using only 10% of the labeled data. This work is the first to reformulate intent classification as an annotation-free, cross-lingual similarity search task.
Resource-constrained edge devices face challenges in accurately understanding user intent from UI interaction traces, while simultaneously ensuring privacy preservation and real-time responsiveness. Method: This paper proposes a two-stage decomposed architecture: (1) generating structured sequential summaries of interaction behaviors, followed by (2) lightweight intent inference based on these summaries. The approach integrates context-aggregated enhancement and task-adaptive fine-tuning to strengthen semantic modeling capabilities of small models. Contribution/Results: Experimental results demonstrate that, under identical privacy guarantees and low-latency constraints, the proposed method achieves higher intent recognition accuracy than state-of-the-art large multimodal language models. It establishes an efficient, privacy-aware, and real-time interaction understanding paradigm for on-device intelligent agents.
This study addresses the limitations of existing query intent classification approaches, which predominantly rely on system logs and neglect user task context, thereby falling short in supporting the nuanced understanding required for complex tasks in the era of large language models. Drawing on grounded theory, the authors conduct qualitative interviews with airport information officers and integrate task analysis with intent modeling to develop the first query intent classification framework that explicitly incorporates task context. Moving beyond traditional models that treat information needs in isolation, the proposed framework introduces a three-dimensional taxonomy of task-oriented information requests—encompassing rules, resources, and constraints—offering a novel paradigm for task-driven retrieval and interaction with large language models.
New intent discovery aims to automatically identify unknown intent categories from unlabeled user utterances to enable continual expansion of dialogue systems; however, existing approaches heavily rely on large-scale manual annotations, suffer from severe noise in pseudo-labels, and exhibit low clustering efficiency. This paper proposes a multi-task pretraining framework for unsupervised new intent discovery: (1) collaborative pretraining leveraging both external labeled data and massive unlabeled corpora; (2) a novel contrastive loss function that exploits self-supervised signals from unlabeled data to enhance discriminability of semantic representations; and (3) integration with an improved unsupervised/semi-supervised clustering algorithm for high-quality intent discovery. Evaluated on three standard benchmarks, our method significantly outperforms state-of-the-art approaches, achieving substantial gains in accuracy and robustness—particularly under zero-shot and few-shot settings.
To address two key limitations of large language models (LLMs) in multi-intent spoken language understanding (SLU)—word-level alignment bias arising from autoregressive generation and insufficient modeling of semantic-level intent relationships—this paper proposes the EN-LLM and ENSI-LLM architectures. Methodologically, it introduces: (1) a restructured entity-slot modeling paradigm; (2) the Sub-Intent Instruction (SII) mechanism, the first of its kind to explicitly guide fine-grained intent decomposition; (3) synthetic benchmarks LM-MixATIS and LM-MixSNIPS; and (4) two novel fine-grained evaluation metrics—Entity Slot Accuracy (ESA) and Combined Semantic Accuracy (CSA). Experiments demonstrate state-of-the-art performance on multi-intent SLU tasks, with significant improvements in robustness to complex intent compositions and distributional shifts, as well as enhanced generalization capability.
This work addresses the performance degradation of existing out-of-scope intent detection methods as the number of known intent classes increases, as well as the high computational cost and deployment challenges associated with large language model–based embeddings. To overcome these limitations, the authors propose a lightweight and efficient approach that leverages MiniLM (all-MiniLM-L6-v2) embeddings within a one-class classification framework augmented with a multi-cluster boundary learning mechanism. This mechanism explicitly models the semantic cluster structure of in-scope training utterances to better characterize the known intent distribution, thereby improving the rejection of out-of-domain samples. The method achieves state-of-the-art performance on CLINC150, StackOverflow, and Banking77 benchmarks. Ablation studies confirm the synergistic benefits of combining MiniLM embeddings with multi-cluster boundaries, demonstrating significantly enhanced out-of-scope detection while maintaining low computational overhead.
This study addresses the challenge of reliably maintaining user goal consistency across multiple models, languages, and prompting frameworks. The authors propose a protocol-like communication layer grounded in structured intent representations—specifically the 5W3H schema—and systematically evaluate its cross-lingual and cross-domain alignment efficacy on Claude, GPT-4o, and Gemini 2.5 Pro. Leveraging both automated evaluation with DeepSeek-V3 and a user study involving 50 participants, the results demonstrate that structured prompting substantially reduces cross-lingual goal drift, lowering the standard deviation of alignment scores from 0.470 to 0.020. The findings further reveal a weak-model compensation effect and the critical role of dimensional decomposition. Notably, Gemini exhibits a +1.006 improvement in goal alignment, accompanied by a 60% reduction in interaction turns and an increase in user satisfaction from 3.16 to 4.04.
This study addresses the problem of efficient and accurate user intent classification in large language models to facilitate downstream domain-specific model routing. For the first time, it systematically compares training-free strategies—such as lightweight methods based on internal representation statistics—with training-based approaches, including linear probes and MLP classifiers, evaluating their performance across varying task difficulty, mixed-intent prompts, and adversarial inputs. The results reveal that both paradigms achieve performance saturation on simple tasks; however, training-based methods excel in fine-grained classification (e.g., distinguishing Java from Python), whereas training-free methods demonstrate superior robustness to mixed and adversarial prompts. These findings highlight fundamental differences between the two approaches in terms of accuracy, robustness, and failure modes.
This work investigates the degradation of recurrent neural network (RNN) performance in intent detection under class-imbalanced data, a phenomenon whose underlying mechanism remains poorly understood. Leveraging dynamical systems theory, we model the evolution of sentence representations in the RNN hidden state space as trajectories and reveal that the network performs intent classification through clustering on low-dimensional manifolds. We propose a novel framework—decoupling geometric separation from readout alignment—and, for the first time, elucidate from a dynamical systems perspective how class imbalance disrupts the geometric structure of hidden states. On the SNIPS dataset, clear clustering structures emerge, whereas in ATIS, clusters corresponding to low-frequency intents significantly deteriorate, thereby uncovering how data distribution shapes the network’s computational solution.
This work addresses the degradation of classification performance in low-resource intent recognition tasks caused by semantically ambiguous samples generated by large language models. To mitigate this issue, the authors propose the DDAIR framework, which introduces a semantic similarity mechanism based on Sentence Transformers to detect and filter cross-category ambiguous samples. Furthermore, an iterative regeneration strategy is employed to enhance the class discriminability of the synthesized data. By explicitly resolving intent boundary ambiguity, the proposed approach significantly improves intent recognition accuracy under low-resource conditions.