Score
Designs, builds, and analyzes techniques, pipelines, and evaluation methods that anchor language-model or agent outputs to external sources or structured data so generated responses are supported, attributable, and verifiable. This includes retrieval and real-time grounding, structured-data grounding, strategies for grounded generation and agent grounding, creation and curation of ground-truth and ground-truth standards, and development of groundedness metrics, attribution methods, and grounding evaluation frameworks.
This study addresses key challenges in deeply integrating large language models (LLMs) with structured knowledge systems—particularly knowledge graphs—including knowledge accuracy, dynamic updating, trustworthy reasoning, and ethical governance. Methodologically, it introduces the first multidimensional evaluation framework for LLM–knowledge base integration, formalizing three core benefits: data contextualization, precision enhancement, and knowledge utilization efficiency, while identifying critical gaps in scalability, real-time knowledge updating, and neuro-symbolic synergy. The approach unifies knowledge graph embedding, retrieval-augmented generation (RAG), prompt engineering, knowledge distillation, and explainability analysis to balance logical rigor with generative flexibility. Drawing on a systematic review of 200+ scholarly works, the study establishes a taxonomy and derives six actionable, industry-ready implementation guidelines. Results provide reusable integration paradigms and risk-mitigation pathways for high-stakes domains including finance, healthcare, and public administration.
To address the challenges of tracing provenance and ensuring credibility of large language model (LLM)–generated content, this paper proposes the first holistic four-dimensional provenance framework integrating both model- and data-centric perspectives: model origin identification, architectural and mechanistic analysis, training data attribution, and external information verification. We introduce a novel “prior–posterior” dual-paradigm classification system and unify techniques including model fingerprinting, response-level verification, and traceability-aware embedding to support both proactive and reactive reasoning. The framework systematically consolidates fragmented provenance research efforts, significantly enhancing the explainability, verifiability, and transparency of AI-generated content. It establishes a theoretical foundation and scalable technical methodology for detecting AI-generated content (AIGC), identifying model identities, and ensuring information reliability.
This study addresses the limitations of prevailing binary support/refutation frameworks in evaluating AI-generated text, which fail to capture the nuanced semantic relationships between generated content and source documents. Moving beyond conventional groundedness paradigms, the work proposes a reader-centered, fine-grained taxonomy of evidential relations by integrating insights from linguistics and philosophy of language, encompassing diverse linkage types such as syntactic rephrasing and inferential strategies. Through theoretical analysis, a human annotation protocol, and benchmark evaluations, the authors systematically demonstrate the feasibility and efficacy of this framework. The resulting approach offers a more transparent and interpretable provenance mechanism for AI outputs, establishing both theoretical foundations and practical pathways for fine-grained evaluation and explainable interfaces in natural language generation systems.
This paper addresses the limitation of large language models (LLMs) in real-world information-seeking tasks requiring integration of multi-source evidence to verify hypotheses, introducing for the first time the “integrative grounding” task and evaluation framework. Methodologically, it constructs a cross-domain benchmark, proposes a premise-induction–driven retrieval planning strategy, and incorporates a zero-shot self-reflection mechanism to enhance evidence verification quality. Key contributions include: (1) revealing that LLMs frequently generate knowledge hallucinations under incomplete information; (2) demonstrating that premise induction significantly outperforms undirected retrieval in suppressing noise and improving evidence relevance; and (3) finding that while LLMs exhibit robustness to redundant evidence, external constraints are necessary to prevent overreliance on internal parametric knowledge. Empirical results validate the framework’s effectiveness in improving multi-evidence reasoning capabilities.
This work addresses systematic limitations in existing large language model evaluation frameworks—particularly their inadequacies in distributional coverage, temporal dynamics, scope, and procedural fidelity—which hinder effective assessment of embodied agents’ long-term reasoning and behavior and exacerbate reward hacking in reinforcement learning from human feedback (RLHF). To overcome these issues, the authors propose the Grounded Continuous Evaluation (GCE) framework, which introduces a simulation-based ISOPro system that replaces learned reward models with deterministic ground-truth verifiers, thereby structurally eliminating reward hacking. GCE enables LoRA weight updates on the CPU, drastically lowering hardware requirements, and pioneers a continuous evaluation paradigm that embeds assessment directly into training, implicitly inducing curriculum formation without manual design. Using only 0.216% trainable parameters, GCE achieves threefold higher accuracy than zero-shot baselines on resource scheduling tasks and demonstrates, for the first time on consumer-grade hardware, emergent capabilities contingent on continuous evaluation.
Large language models (LLMs) frequently exhibit intent drift, contextual incoherence, and factual hallucinations in dialogue, undermining their reliability in real-world applications. To address this, we conduct a rapid systematic review guided by the PRISMA framework and the PICO strategy. This work introduces— for the first time—a taxonomy of dialogue alignment techniques structured along the LLM lifecycle: inference-time, post-training, and reinforcement learning stages. We particularly highlight inference-time interventions—including prompt engineering, self-verification, and retrieval-augmented generation—which improve intent consistency, contextual groundedness, and hallucination suppression *without* model retraining. Empirical findings demonstrate that these methods offer high efficiency, practical deployability, and flexibility across diverse deployment scenarios. Our taxonomy and analysis thus provide both a theoretically grounded framework and an actionable technical pathway for enhancing dialogue reliability in production LLM systems.
Enterprise AI systems struggle to gain trust due to hallucinations—confident yet incorrect outputs from large language models—a risk that existing approaches fail to eliminate. This work proposes HALO, a novel architecture that treats hallucination as a manageable system failure mode and enforces a “zero-hallucination” guarantee through six coordinated defense layers: retrieval-constrained generation, deterministic execution constraints, multi-signal validation (integrating LLM-based discriminators with source document alignment), calibrated refusal, end-to-end traceability, and a continuous monitoring feedback loop. Evaluated on insurance claim information extraction, HALO delivers high-fidelity outputs, effectively blocks hallucinations, and provides early warnings of system drift, substantially enhancing the trustworthiness of enterprise AI deployments.
This work addresses the limitations of existing large language model (LLM)-based methods for synthesizing agent interaction data, which often suffer from insufficient realism and diversity, thereby hindering their applicability to high-fidelity, long-horizon complex tasks. The paper proposes GAIS, a novel framework that uniquely integrates real-world Model Context Protocol (MCP) server environments with structured task planning. By anchoring interactions in authentic protocol contexts and incorporating logical dependency constraints, structure-guided planning, and adversarial strategy generation, GAIS produces diverse and complex tasks while mitigating the data bias inherent in purely LLM-generated approaches. Empirical evaluations on BFCL, τ²-Bench, and ACEBench demonstrate that GAIS significantly outperforms current methods, enabling base models to achieve performance comparable to—or even exceeding—that of officially instruction-tuned variants, all while offering superior data efficiency and scalability.
Traditional industries rely on unstructured long-form documents to store high-value information, yet they lack high-quality, scalable training data for text-to-JSON structuring. This work proposes STAGE, a method that achieves the first end-to-end synthetic data generation pipeline grounded in spreadsheet source data: leveraging large language models to produce semantically coherent narrative reports alongside their corresponding JSON representations, and employing value-level verification against the original source data to ensure factual fidelity. The STAGE-Eval dataset constructed via this approach substantially enhances model performance—on Qwen3-4B, exact match accuracy improves from 31.37% to 74.27%, and value-level accuracy rises from 45.46% to 90.69%.