query bibliographic databases

Design and implement structured query formulations and search strategies for bibliographic databases, including Boolean, metadata- and field-specific filters, iterative query refinement, and scalable retrieval workflows. Use these queries and strategies to retrieve, filter, deduplicate, and resolve bibliographic records at scale and to systematically identify relevant literature and domain datasets.

querybibliographicdatabases

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Database-Augmented Query Representation for Information Retrieval

Jun 23, 2024
SJ
Soyeong Jeong
🏛️ Korea Advanced Institute of Science and Technology

To address the low retrieval accuracy caused by short user queries, this paper proposes DAQu, a database-augmented query representation framework that dynamically expands query semantics using relational database metadata across multiple interlinked tables. Methodologically, it introduces a novel graph-structured unordered set encoding strategy to model hierarchical cross-table metadata relationships and integrate high-dimensional heterogeneous features; it jointly optimizes graph neural networks, metadata modeling, set encoding, query expansion, and dense retrieval. On multi-scenario retrieval tasks, DAQu achieves a 12.7% improvement in Recall@10 over state-of-the-art query enhancement methods, demonstrating the substantial benefit of leveraging structured database knowledge for query representation. The core contribution lies in formulating database metadata as a learnable, graph-structured prior and enabling end-to-end semantic-enhanced retrieval.

Addressing short query challenge in information retrievalAugmenting queries with metadata from relational databasesEncoding unordered metadata features via graph-based strategy

Existing databases struggle to efficiently support hybrid queries over structured and unstructured (e.g., vector) data, resulting in poor performance for joint semantic retrieval and SQL execution. This paper introduces the first full-stack native hybrid query engine. Our approach addresses this challenge through three core innovations: (1) semantic-aware query classification and dynamic physical plan optimization; (2) customized physical operators that eliminate redundant computation; and (3) a JIT-compilation-based execution framework tailored for vector–relational hybrid workloads, integrating approximate nearest-neighbor indexing with a unified hybrid query optimizer. Evaluated on real-world datasets, our engine achieves end-to-end query speedups of 13%–7500× over state-of-the-art systems. These gains significantly enhance efficiency in multimodal recommendation and analytical scenarios requiring tight coupling of semantic and relational operations.

Database SystemsMixed Regular and Irregular DataQuery Efficiency

Exploring Multi-Table Retrieval Through Iterative Search

Nov 17, 2025
AB
Allaa Boutaleb
🏛️ Sorbonne Université | CNRS | LIP6

This work addresses the challenge of cross-table retrieval and information composition in open-domain question answering. We propose an iterative multi-table retrieval framework that jointly optimizes semantic relevance, query coverage, and structural connectability. To our knowledge, this is the first approach to formulate multi-table retrieval as a greedy iterative search process, incorporating a lightweight joint-aware algorithm that dynamically evaluates semantic matching, coverage completeness, and inter-table joinability at each step. Evaluated on five mainstream NL2SQL benchmarks, our method achieves retrieval accuracy comparable to exact MIP solvers while accelerating inference by 4–400×, significantly outperforming conventional single-objective heuristic methods. Our core contributions are: (i) the first iterative retrieval paradigm that jointly optimizes semantic relevance, coverage, and structural connectability; and (ii) an efficient, interpretable, and scalable solution for multi-table joint retrieval.

Balancing computational complexity with joinability optimization in table retrievalDeveloping scalable iterative methods for multi-table question answering systemsRetrieving semantically relevant and structurally coherent multi-table data

To address scalability bottlenecks in Text-to-SQL for enterprise-scale databases, this paper proposes a domain-agnostic, retrieval-augmented schema linking framework. Methodologically, it introduces a modular semantic indexing architecture that decouples database schemas and metadata into fine-grained semantic units for independent vectorization; further, it designs a table-level prioritized identification mechanism coupled with column-level contextual fusion, integrating chunked indexing, semantic retrieval, recall re-ranking, and context budget control. Crucially, the approach fully leverages intrinsic semantic cues from raw metadata without requiring domain-specific fine-tuning, enabling high-precision schema matching. Experimental results demonstrate that our method consistently outperforms mainstream baselines across heterogeneous, multi-source data catalogs—achieving both high recall and high accuracy. Its plug-and-play compatibility facilitates seamless enterprise deployment.

Leveraging database metadata for semantic context in retrievalMaintaining accuracy without domain-specific fine-tuningScaling text-to-SQL to enterprise databases with massive schemas

This study addresses the challenge of efficiently supporting direct access to database query results by rank position, particularly under diverse query types and sorting strategies. To this end, we develop a scalable direct-access system that integrates multiple high-performance algorithms and conduct a systematic experimental evaluation on mainstream database platforms. Our work presents the first empirical analysis of direct-access algorithms across a broad spectrum of query and ranking scenarios, bridging a critical gap between theoretical proposals and practical deployment. The experiments reveal significant performance variations among databases in handling direct-access workloads, validate the real-world effectiveness—and limitations—of existing algorithms, and provide empirical insights to guide future optimizations.

database systemsdirect accesspractical performance

Latest Papers

What's happening recently
View more

This work addresses the overreliance on extremely large language models in scientific knowledge discovery, which hinders reproducibility and accessibility. The authors propose a lightweight retrieval-augmented framework featuring a task-aware retrieval routing mechanism that dynamically selects appropriate strategies by integrating full-text content with structured metadata. Coupled with a small instruction-tuned language model, this approach generates citation-grounded responses. Experimental results demonstrate that the method substantially enhances the performance of small models across diverse tasks—including scholarly question answering, biomedical question answering, and text summarization—showcasing that well-designed retrieval mechanisms can effectively compensate for limited model capacity. The findings further reveal a complementary relationship between retrieval design and model scale, offering a novel paradigm for building efficient, reproducible academic AI assistants.

accessibilitylarge language modelsmodel scale

This work addresses the unclear efficacy of query decomposition across different stages in multi-condition retrieval, where early-stage decomposition often leads to semantic dilution and degraded performance. Through empirical analysis, the study reveals that the effectiveness of query decomposition is highly stage-dependent and proposes the first stage-aware query decomposition framework. The approach preserves the full query during initial retrieval to maintain global semantics while introducing subqueries in the reranking stage to enable fine-grained constraint matching. Evaluated on the MultiConIR and SSRB benchmarks, this method significantly enhances the ranking performance of diverse retrieval and reranking models on compositional queries, effectively balancing semantic completeness with local constraint verification.

multi-condition retrievalquery decompositionreranking

This work addresses the limitation of existing deep research agents that overlook structured web fields—such as titles, sections, and metadata—during the search-and-retrieve process, leading to redundant retrieval and excessive contextual noise. To overcome this, the authors propose SIEVE, a novel framework that introduces fielded Boolean queries (BQL) into the research agent paradigm for the first time. SIEVE implements a three-stage pipeline: search, inspect, and retrieve. It first filters candidate pages using document-level fields, then presents results as structured cards for selective inspection, and finally retrieves only the content of chosen sections. By leveraging web structure at a fine-grained level, SIEVE achieves higher accuracy than state-of-the-art baselines across three question-answering benchmarks while reducing context token consumption by 20.7%–50.6%, demonstrating both efficiency and broad applicability.

Boolean retrievaldeep-research agentsdocument structure

Scientific publishing agents often generate field-level errors in BibTeX citations due to overreliance on parametric memory. To address this, this work constructs a benchmark of 931 papers spanning four disciplines and three citation tiers, and for the first time disentangles model memory from retrieval dependence under realistic search conditions. The authors propose a two-stage citation correction framework based on co-occurrence patterns of field-level errors. By integrating large language models—including GPT-5, Claude Sonnet-4.6, and Gemini-3 Flash—with deterministic retrieval tools such as Zotero and CrossRef, the method improves field-level accuracy from 83.6% to 91.5% and full-entry correctness from 50.9% to 78.3%, with a fallback rate of only 0.8%.

BibTeXcitation hallucinationlarge language models

Hot Scholars

CL

Carsten Lutz

Professor of Computer Science, University of Leipzig
Knowledge RepresentationArtificial IntelligenceLogic in Computer ScienceTheoretical Computer Science
TG

Tirthankar Ghosal

Oak Ridge National Laboratory
Natural Language ProcessingMachine LearningArtificial IntelligenceInformation Extraction
ZT

Zeerak Talat

University of Edinburgh
NLPOnline AbuseHate SpeechSTS
AA

Arif Ali Khan

Associate Professor (tenure track), M3S Research Unit, University of Oulu, Finland
Quantum Software EngineeringGlobal Software DevelopmentSoftware Process Improvement