Score
Designs and implements retrieval-augmented generation systems that use a structured taxonomy to constrain, filter, and contextualize retrieved evidence; builds taxonomy-based retrieval indexes and query pipelines that restrict searches to specific taxonomy nodes or versions. Evaluates and refines model outputs by anchoring generated suggestions to taxonomy entries and evidence to improve fidelity and precision.
This work addresses the lack of a systematic taxonomy for retrieval-augmented generation (RAG) applications. We propose the first comprehensive, lifecycle-spanning classification framework for RAG applications. Methodologically, we introduce a novel four-stage iterative construction paradigm—comprising multi-round expert collaboration, systematic literature review, dimensional abstraction, and empirical validation—thereby filling a critical gap in classification research beyond the ACL community. The framework comprises five meta-dimensions and sixteen fine-grained dimensions, balancing structural clarity with extensibility. Empirical validation across education, healthcare, and legal domains demonstrates its effectiveness in supporting design decisions, technical evaluation, and cross-domain understanding of RAG applications. By providing a foundational taxonomic infrastructure, this work advances the engineering-oriented deployment and standardization of RAG systems.
To address core limitations of large language models (LLMs)—including hallucination, knowledge obsolescence, and poor domain adaptability—this work systematically advances the Retrieval-Augmented Structured (RAS) generation paradigm. We propose a multi-granularity knowledge acquisition mechanism integrating sparse, dense, and hybrid retrieval, coupled with text structuralization, taxonomy construction, knowledge embedding, and prompt-driven reasoning to enable efficient external knowledge retrieval, semantic alignment, and controllable integration. Crucially, we deeply embed structured modeling into the augmentation pipeline, enhancing factual accuracy, temporal freshness, and domain-specific competence of generated outputs. Our contributions include: (1) a unified methodological framework for RAS generation; (2) principled pathways toward multimodal, cross-lingual, and interactive augmented generation; and (3) empirically validated improvements in reliability and specialization across diverse domains. This work establishes foundational design principles and future research directions for next-generation RAS systems.
To address error propagation and answer bias arising from unidirectional retrieval-then-reasoning in multi-hop question answering, this paper proposes RetroRAG, the first framework introducing backtracking-style reasoning. Its core is an evidence backtracking mechanism: inferring entity-centric queries to dynamically revise retrieved evidence and reconstruct reasoning paths, enabling iterative refinement and dynamic reorganization of trustworthy evidence through coordinated multi-round retrieval-generation-evaluation cycles. This establishes a closed-loop “evidence curation–discovery–verification” process, substantially enhancing robustness and interpretability for complex reasoning. On mainstream multi-hop QA benchmarks, RetroRAG consistently outperforms existing RAG methods, achieving significant gains in answer accuracy—particularly under challenging conditions involving long reasoning chains and noisy evidence.
Traditional RAG systems often suffer from contextual redundancy, low information density, and fragile reasoning in multi-hop question answering due to unstructured retrieval and single-pass generation. This work proposes a structured reasoning framework that eschews explicit graph construction by representing queries and documents as relational triples. It employs a lightweight two-stage classification mechanism to constrain entity semantics, decomposes complex questions into ordered sub-queries, and performs stepwise evidence selection by jointly leveraging semantic similarity and structural consistency. An explicit entity binding table is introduced to resolve intermediate variables and disambiguate entities. The approach outperforms strong baselines by up to 14% across multiple multi-hop QA benchmarks while yielding more interpretable evidence tracing and trustworthy reasoning trajectories.
Scientific document retrieval faces significant challenges due to the scarcity of domain-specific labeled data and the highly specialized nature of technical terminology, which often leads existing methods to suffer from conceptual redundancy or insufficient coverage. To address these limitations, this work proposes an academic concept indexing framework that integrates a structured scholarly taxonomy with large language models to extract and organize key concepts. The framework introduces two novel mechanisms: Concept-Coverage-aware Query Generation (CCQGen) and Concept-Focused Context Expansion (CCExpand), which jointly enhance the retrieval system’s capacity to understand and match scientific semantics. Experimental results demonstrate that the proposed approach substantially improves query quality, concept alignment, and overall retrieval effectiveness, outperforming current state-of-the-art methods on scientific document retrieval benchmarks.
This study addresses the limitations of semantic similarity–based retrieval in structured, highly repetitive regulatory texts, where linguistic overlap often obscures meaningful content distinctions and undermines the effectiveness of retrieval-augmented generation (RAG). To mitigate this issue, the work systematically investigates metadata-aware retrieval strategies, proposing and evaluating fusion approaches such as unified embedding and prefix concatenation. The findings demonstrate that incorporating metadata enhances intra-document cohesion and reduces inter-document ambiguity, thereby improving retrieval performance. Evaluated on a newly curated benchmark dataset, RAGMATE-10K, both the unified embedding and prefix-based methods significantly outperform pure text baselines across multiple question types and evaluation metrics. Notably, the unified embedding approach achieves superior performance while maintaining greater maintainability.
This work addresses the challenge in LongEval-RAG tasks where responses must be strictly grounded in a given set of candidate documents. To this end, the authors propose a candidate-constrained retrieval-augmented generation (RAG) system that integrates rule-based chunking, query expansion, pseudo-relevance feedback, reciprocal rank fusion, MiniLM sentence-level reranking, and citation-aware evidence aggregation, complemented by deterministic provenance tracing and a neural sentence selection mechanism. Experimental results demonstrate that the proposed rule-MiniLM variant significantly outperforms baselines across multiple metrics—including BERTScore, retrieval precision, information point coverage, and human evaluation—thereby validating the effectiveness of combining rule-based chunking with neural sentence selection. The study further underscores the critical role of multi-metric evaluation in diagnosing and advancing RAG system performance.
This work addresses the "retrieval readiness gap"—a performance bottleneck in large-scale classification systems when inputs consist solely of indirect evidence (e.g., table cells) and suffer from semantic ambiguity. To overcome this, the authors propose a factorized hypothesis search mechanism that decomposes semantic interpretation into composable, named dimensional hypotheses. By leveraging structured query generation, parallel multi-hypothesis retrieval, and dimension-level candidate validation, the approach circumvents reliance on free-form text generation. Evaluated on financial taxonomy labeling and the CodiEsp clinical coding task, the method significantly outperforms existing non-oracle approaches, achieving consistent improvements in Recall@1, Mean Reciprocal Rank (MRR), and final accuracy, thereby demonstrating its effectiveness and robustness.
To address insufficient evidence coverage and low answer accuracy in Retrieval-Augmented Generation (RAG) for government document question answering within legal and regulatory domains, this paper proposes two synergistic optimization strategies. First, a One-SHOT retrieval method with adaptive token budgeting improves recall of critical information chunks. Second, an iterative retrieval framework built upon a Reasoning Agentic RAG architecture integrates dynamic query generation, progressive context refinement, and result evaluation with feedback—effectively mitigating query drift and retrieval inertia. Experimental results demonstrate substantial improvements: +28.6% in evidence coverage and +14.3 BLEU points in answer accuracy. These advances establish a novel paradigm for high-precision, interpretable legal intelligent question answering.