Intent Classification on Low-Resource Languages with Query Similarity Search

📅 2025-05-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of ambiguous intent definitions, high annotation costs, and data scarcity in low-resource language intent classification, this paper proposes a zero-shot retrieval-based intent identification method. Instead of relying on explicit intent labels, the approach models each intent as a set of historically annotated queries and performs intent inference via dense retrieval—matching an input query to its most semantically similar labeled queries in a shared latent space. The method employs a multilingual semantic encoder to produce cross-lingually aligned query embeddings, enhances semantic consistency through contrastive learning, and accelerates retrieval using approximate nearest neighbor (ANN) search. Evaluated on eight low-resource languages, it achieves an average F1 score of 72.4%, outperforming zero-shot fine-tuning baselines by 18.6 percentage points; remarkably, it attains near fully supervised performance using only 10% of the labeled data. This work is the first to reformulate intent classification as an annotation-free, cross-lingual similarity search task.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Representation, semantic annotation, enhancement, enrichments, access and/or integration of a variety of data on the Web
📝 Abstract
Intent classification is an important component of a functional Information Retrieval ecosystem. Many current approaches to intent classification, typically framed as a classification problem, can be problematic as intents are often hard to define and thus data can be difficult and expensive to annotate. The problem is exacerbated when we need to extend the intent classification system to support multiple and in particular low-resource languages. To address this, we propose casting intent classification as a query similarity search problem - we use previous example queries to define an intent, and a query similarity method to classify an incoming query based on the labels of its most similar queries in latent space. With the proposed approach, we are able to achieve reasonable intent classification performance for queries in low-resource languages in a zero-shot setting.
Problem

Research questions and friction points this paper is trying to address.

Intent classification in low-resource languages
High cost of annotated data for intent definition
Zero-shot query similarity for intent classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Query similarity search for intent classification
Zero-shot performance in low-resource languages
Latent space comparison of query similarity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arjun Bhalla
Bloomberg L.P.
Q
Qi Huang
Bloomberg L.P.