query understanding

Designs and implements components that interpret and transform user queries by extracting intent, entities, and structural meaning, performing parsing, normalization, disambiguation, and semantic representation to make queries actionable for retrieval, execution, or QA systems. Analyzes query logs and interaction signals to model ambiguity, generate rewrites or suggestions, and measure how accurately query interpretations support downstream tasks.

queryunderstanding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

The Case for Intent-Based Query Rewriting

Nov 25, 2025
GL
Gianna Lisa Nicolai
🏛️ RPTU Kaiserslautern-Landau

This paper addresses the failure of conventional equivalence-based query rewriting when original data tables are inaccessible due to access control policies, privacy constraints, or prohibitive retrieval costs. To overcome this, we propose INQURE, a semantic intent-preserving query rewriting framework. Unlike traditional approaches relying on syntactic equivalence and query plan optimization, INQURE introduces, for the first time, large language model (LLM)-driven intent understanding and cross-table reconstruction—enabling semantically consistent rewriting across structurally heterogeneous and non-aligned schemas. The system incorporates pre-filtering of candidate tables, pruning heuristics, and learned ranking to form an end-to-end rewriting pipeline. Evaluated on a benchmark spanning 900+ real-world database schemas, INQURE demonstrates superior rewriting quality and practical utility. A user study further confirms its effective trade-off between execution feasibility and fidelity of analytical insights.

Developing intent-based query rewriting using large language modelsEnabling data access despite access control, privacy, or cost constraintsRewriting queries to preserve insights while altering structure and syntax

The Interpretability Analysis of the Model Can Bring Improvements to the Text-to-SQL Task

Aug 12, 2025
CZ
Cong Zhang
🏛️ China Life Insurance Company Limited

Text-to-SQL models suffer from strong dependence on condition-column data distributions and human annotations in real-world settings, coupled with poor generalization in semantic parsing of WHERE clauses. To address these issues, this paper proposes a condition-augmentation method grounded in interpretability analysis and execution-guided learning. Our approach employs three core mechanisms—filtering-and-adjustment, logical-association refinement, and multi-model fusion—operating without additional labeled data, thereby substantially reducing model sensitivity to the value distribution of condition columns. Evaluated on the WikiSQL single-table benchmark, the method achieves significant improvements in overall execution accuracy. Notably, it sets a new state-of-the-art (SOTA) on the subtask of condition-value prediction within WHERE clauses, demonstrating enhanced fundamental semantic understanding and superior cross-instance generalization capability.

Enhancing text-to-SQL model generalization through interpretability analysisImproving WHERE clause semantic parsing accuracy in database queriesMinimizing dependency on condition column data and manual labels

Query Understanding in LLM-based Conversational Information Seeking

Apr 08, 2025
YY
Yifei Yuan
🏛️ University of Copenhagen | University of Amsterdam | Singapore Management University

Conversational information retrieval (CIS) faces significant challenges due to high user intent volatility, strong query ambiguity, and deep contextual dependencies. Method: This paper proposes an LLM-driven multi-turn query understanding framework featuring a context-aware intent parsing model that integrates instruction tuning with interactive reasoning. It introduces, for the first time, proactive query management and adaptive query reconstruction mechanisms—overcoming limitations of conventional static query modeling. Technically, the framework incorporates LLM-based contextual encoding, dynamically constructed evaluation metrics, and an interpretability verification module. Contribution/Results: Experiments on mainstream CIS benchmarks demonstrate a 12.3% improvement in intent identification F1-score and substantial gains in query rewriting quality. Furthermore, we release the first open-source evaluation protocol specifically designed for LLM-enhanced CIS, establishing foundational standards for systematic, reproducible benchmarking in this emerging domain.

Addressing challenges in LLM integration for conversational searchDeveloping robust evaluation metrics for multi-turn interactionsEnhancing query understanding in LLM-based conversational search systems

QUIDS: Query Intent Generation via Dual Space Modeling

Oct 16, 2024
YW
Yumeng Wang
🏛️ Leiden University | Mohamed bin Zayed University of Artificial Intelligence

To address the challenges of ambiguous user queries and insufficient intent feedback in exploratory search—leading to iterative trial-and-error—we propose, for the first time, the generative query intent description task. Our method jointly models relevant and irrelevant documents to automatically generate precise, interpretable intent descriptions. Technically, we introduce a novel dual-space modeling paradigm: semantic separation of relevant and irrelevant documents in a representation space, and explicit suppression of irrelevant information in a disentangled space. We design a semantic projection encoder, a semantic disentanglement decoder, and an attention-guided generation framework. Experiments on benchmark datasets demonstrate that our approach significantly outperforms existing intent classification, clustering, and query summarization methods. Attention visualization confirms its effectiveness in filtering out distracting topics, yielding more accurate and interpretable intent descriptions.

Addresses vague queries in exploratory search through dual-space modelingGenerates natural language descriptions of search query intentImproves user-system interaction by providing interpretable search feedback

Towards a Unified Query Plan Representation

Aug 14, 2024
JB
Jinsheng Ba
🏛️ National University of Singapore

Database query plan representations are highly fragmented, impeding test method reuse and cross-system analysis. Method: This paper proposes the first database-agnostic unified query plan representation framework, systematically identifying the “operator–attribute–format” trinity as the common structural foundation across execution plans. It abstracts internal plans from nine mainstream databases via cross-database reverse parsing and intermediate representation modeling, yielding an extensible, formally verifiable unified model. Contribution/Results: The framework enables seamless reuse of existing testing methodologies across all nine databases, uncovering 17 previously undetected, database-specific defects. It facilitates rapid adaptation of multi-database visualization tools and supports standardized comparative analysis—including semantic alignment and performance profiling—of query plans across heterogeneous systems.

Enabling cross-system performance comparison and optimization insightsReducing implementation effort for testing and visualization toolsUnifying diverse query plan representations across database systems

Latest Papers

What's happening recently
View more

This work addresses the inefficiency faced by data analysts who must repeatedly submit and integrate multiple related queries to explore salient data patterns. To streamline this process, the paper introduces the ANALYZE operator, which formalizes such exploratory analysis as five auxiliary cube queries, enabling comprehensive 360-degree examination of specific data subsets. Leveraging multi-query optimization (MQO), the authors devise three query merging and execution strategies—Mid-MQO, Min-MQO, and Max-MQO—that significantly improve execution efficiency while preserving result equivalence. Experimental evaluation demonstrates that Mid-MQO consistently delivers the best overall performance across most scenarios, whereas Max-MQO excels when sibling queries are numerous and exhibit high overlap.

ANALYZE operatorcube queryingdata analysis

Existing natural language interfaces to databases lack systematic evaluation frameworks and design theories. This work proposes QUEST, a novel framework that integrates the general-purpose FAR operation schema—Filter, Aggregate, Return—with the W5H semantic dimensions (Who, What, Where, When, Why, How) to enable structured analysis and evaluation of text-to-SQL query semantics. Through semantic annotation and structural parsing, the study validates the universality of the FAR schema across five cross-domain datasets comprising 120,464 queries. The analysis further reveals significant inter-domain disparities in semantic distributions: for instance, medical queries predominantly focus on WHEN and WHO, while WHY and HOW are nearly absent, underscoring a critical challenge for machines in performing deep reasoning over structured data.

natural language interfacesquery understandingsemantic evaluation

Analytical Search

Feb 12, 2026

Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.

analytical searchevidence fusioninformation retrieval

This work addresses the challenges of fragmented query understanding modules in industrial-scale semantic search, which lead to high maintenance costs and unstable performance on long-tail queries. The authors propose Query Illuminator, a unified structured query understanding framework that consolidates multiple tasks into a single small language model (SLM) and enables end-to-end processing through schema-constrained generation. The framework supports high-quality automatic annotation distillation and scalable evaluation under scarce labeled data, while facilitating cross-domain transfer. Deployed in LinkedIn’s job search system, Query Illuminator significantly improves user engagement and reduces operational overhead, achieving efficient inference under stringent latency constraints and limited GPU resources.

fragmented architectureindustrial semantic searchlong-tail queries

Current information retrieval systems are designed with human users in mind and struggle to accommodate search behaviors initiated by autonomous agents, leading to performance degradation and evaluation bias. To address this gap, this work proposes a systematic approach that leverages a multi-agent framework and diverse retrieval pipelines to collect agent-generated queries, retrieved documents, and reasoning traces on established benchmarks such as HotpotQA, Researchy Questions, and MS MARCO. We construct and release the first dataset specifically tailored to agentic search behavior—Agentic Search Queryset (ASQ)—alongside a supporting toolkit. This resource fills a critical void in authentic interaction data for agent-driven retrieval, enables flexible extension to new agents, retrievers, and tasks, and lays the foundation for future research in agentic information retrieval.

agent behavioragentic searchdataset gap

Hot Scholars

LS

Le Sun

Institute of Software, CAS
information_retrievalnatural_language_processing
XH

Xiangnan He

University of Science and Technology of China
RecommendationCausalityBig DataInformation Retrieval
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
ZC

Zhuangbin Chen

Assistant Professor, School of Software Engineering, Sun Yat-sen University
Software EngineeringDistributed SystemsCloud ComputingLLM Systems
AJ

Adam Jatowt

Professor at Univ. of Innsbruck (previously Kyoto Univ.)
question answeringlarge language modelsinformation retrievalRAG