Score
Designs, builds, or analyzes systems that iteratively generate, evaluate, and update symbolic or retrieved knowledge artifacts (rules, candidate hypotheses, retrieved documents, or model outputs) through automated or human-in-the-loop cycles to improve correctness, coverage, and explainability. This work encompasses pipelines for automated rule evolution and patching, interactive rule editing, retrieval-guided rewriting and refinement, student–teacher or revision loops, and staged multi-step reasoning that compare outputs against evidence or expectations and revise artifacts accordingly.
To address the limitations of existing RAG systems in complex reasoning, dynamic retrieval, and multimodal integration within real-world industrial applications, this paper proposes an inference-enhanced intelligent RAG framework. Methodologically, it introduces the first dual-track reasoning taxonomy—System 1 (fast, modular reasoning) and System 2 (slow, autonomous planning)—and establishes the first open-source knowledge-graph-based RAG survey repository. The framework integrates LLM-driven reasoning architectures, standardized tool-use protocols (e.g., ReAct), multi-stage retrieval strategies, and multimodal interfaces. Through a systematic analysis of over 120 state-of-the-art works, we identify seven inference patterns and five solutions to key industrial bottlenecks. Empirical evaluation in production scenarios—including customer service and financial risk control—demonstrates 23%–38% improvements in reasoning accuracy.
This paper investigates whether large language models (LLMs) can transcend information retrieval and instruction following to achieve genuine novel knowledge discovery. Method: Grounded in Peirce’s abductive–deductive–inductive triadic logic, it establishes the first unified analytical framework for LLM-driven hypothesis generation toward AGI, systematically characterizing critical pathways and fundamental bottlenecks in generative knowledge discovery. It proposes a closed-loop “hypothesis generation–application–validation” technical architecture integrating prompt engineering, self-verifying reasoning, rule distillation, and empirical evaluation. Contribution/Results: Synthesizing over 100 state-of-the-art studies, the work identifies key advances—including transferable hypothesis modeling, domain-adaptive validation, and enhanced causal interpretability—while revealing six persistent challenges: weak falsifiability, poor cross-domain generalization, among others. The framework provides both theoretical grounding and methodological foundations for evolving LLMs into scientific innovation engines.
This work proposes the first end-to-end autonomous research system that is haltable, auditable, and supports human–AI collaboration, aiming to automate the entire scientific workflow from topic selection to manuscript writing. Built upon large language models, the system employs a multi-agent architecture integrated with a unified memory mechanism, open academic indexing for verification, executable code generation, and traceable result provenance. It further incorporates a preregistered outcome contract to enforce an evidence-based closed loop. Researchers can intervene and refine the process at any stage. Experimental results demonstrate that the system can independently or collaboratively produce high-quality, verifiable research papers that adhere to publication standards.
To address the low efficiency of knowledge structuring in scientific literature—exacerbated by its exponential growth and heavy reliance on expert curation—this paper proposes a neural-symbolic, human-in-the-loop (HITL) workflow. The method leverages large language models (LLMs) to automatically extract and structure scholarly information, which is then ingested into the Open Research Knowledge Graph (ORKG). A modular architecture enables customizable LLM selection and multi-stage human verification, tightly coupling automation with expert oversight. Its key innovation lies in pioneering a synergistic mechanism between LLMs and symbolic knowledge graphs within the HITL paradigm. Evaluation shows the system achieves a System Usability Scale (SUS) score of 84.17 (A+ level), and reduces per-paper knowledge modeling time from hours to weeks down to an average of 24 minutes and 40 seconds—demonstrating substantial gains in scientific knowledge transformation efficiency.
Manual dependency, error-proneness, and knowledge maintenance difficulties hinder fault diagnosis in power grids. Method: This paper proposes an automated diagnostic framework integrating explicit procedural knowledge and implicit expert expertise. It employs a multi-agent system incorporating: (i) PASTA-formatted fault trees for structured fault representation; (ii) the AlphaEvolve module for reasoning optimization; (iii) a human-in-the-loop verification interface; and (iv) n8n-based executable workflow synthesis. Crucially, it introduces a novel human-feedback closed-loop mechanism to jointly model regulatory logic and expert experience within executable workflows and enable iterative refinement. Results: Evaluated on a transformer fault dataset, the framework achieves 100% topological consistency and high semantic fidelity. It substantially reduces expert workload and—critically—demonstrates, for the first time, the feasibility and effectiveness of end-to-end automated fault diagnosis.
This work addresses the limitations of existing AI agent workflows, which rely on implicit dialogue states and struggle to ensure stability of intermediate artifacts, isolate irrelevant updates, and propagate changes precisely. To overcome these challenges, the paper proposes modeling AI-native workflows as directed acyclic graphs (DAGs) and introduces the concept of execution lineage. By leveraging explicit dependency tracking, identity-based identification of intermediate artifacts, and an identity-aware replay mechanism, this approach achieves deterministic computation graphs in AI agents for the first time. The method guarantees precise change propagation, zero contamination across unrelated branches, and preserves both upstream stability and cross-artifact consistency. Evaluated on a policy memo updating task, DAG-based replay attains 100% fidelity in final outputs, substantially outperforming iterative baseline approaches.
This work addresses the “reasoning gap” commonly observed in large language models during dynamic knowledge updating—where models retain edited facts but fail to correctly apply them in multi-step reasoning. To bridge this gap, the authors propose MCircKE, a novel framework that introduces mechanistic causal circuit analysis into knowledge editing for the first time. By identifying and precisely adjusting the parameters of causal circuits responsible for specific reasoning tasks, MCircKE jointly localizes factual storage and reasoning pathways, enabling their coordinated editing. The method implements a “map-and-adapt” editing pipeline and achieves substantial improvements on the MQuAKE-3K benchmark in multi-hop reasoning scenarios, effectively overcoming the limitations of conventional approaches that modify isolated facts without considering their downstream inferential use.
Scientific hypothesis generation requires modeling the dynamic evolution of knowledge rather than relying on static snapshots of the literature. This work proposes a Continuous Knowledge Metabolism (CKM) framework that incrementally processes scientific publications through a sliding time window to construct a structured knowledge base, enabling hypothesis generation grounded in trajectories of knowledge evolution—such as novelty, corroboration, and contradiction. Experimental results demonstrate that the lightweight variant, CKM-Lite, significantly outperforms batch-processing baselines in hypothesis hit rate, output volume, and alignment quality while reducing token consumption by 92%. The full-fledged CKM-Full further uncovers a critical trade-off between hypothesis quality and coverage and highlights the pivotal role of domain stability in determining hypothesis success.
Existing datasets of scientific ideation trajectories struggle to comprehensively capture the full research process—from literature exploration and tool utilization to the evolution of intermediate artifacts and final proposals. This work proposes a reverse-to-forward synthesis mechanism that emulates the uncertainty, evidence integration, and phased convergence characteristic of real scientific inquiry through a Generator–Advisor architecture. By leveraging action–observation–editing sequence modeling, context-aware verification, and process-level supervision, the approach generates multi-turn trajectories aligned with authentic research practices, starting from high-quality papers and proposals. The study yields the first trajectory dataset spanning the complete scientific workflow and establishes a generalizable paradigm for synthesizing process-supervised data for scientific agents.
This work addresses the semantic gap between natural language specifications and RTL designs, which often leads to SystemVerilog assertions containing syntactic errors or semantic inaccuracies that hinder formal verification. To bridge this gap, the authors propose a knowledge graph–based multi-agent collaborative framework that unifies specifications, RTL code, and verification feedback into a structured intermediate representation for the first time. This enables traceable, design-anchored contextual modeling and supports a closed-loop assertion refinement process through a triple iterative optimization mechanism—comprising syntax repair, counterexample-guided correction, and coverage-driven enhancement. Evaluated on seven benchmark designs, the generated assertions are all compilable with low syntax-repair overhead and achieve formal verification coverage ranging from 78.5% to 99.4%.
This work addresses the challenges of maintaining documentation in large codebases—namely, the lack of semantic structure in existing tool-generated content and difficulties in tracking changes—by proposing Repository Knowledge Graphs (RepoKG). RepoKG introduces a three-stage pipeline comprising code entity relation extraction, functional module clustering, and agent-driven documentation generation, establishing knowledge graphs as the semantic foundation for the entire documentation lifecycle. It incorporates modular hierarchical organization and a bidirectional semantic influence propagation mechanism to enable structured, cross-referable documentation with efficient incremental updates. Evaluated across 24 multilingual repositories, RepoKG improves API coverage by 32.5% and completeness by 10.4%, while accelerating generation by 3× and reducing token consumption by 85%. For incremental updates, it cuts update time by 73%, lowers token usage by 77%, and increases update recall by 10.2%.
This study addresses the underexplored nature of rules in AI-powered integrated development environments (IDEs) as an emerging class of software artifacts, whose taxonomy, evolution patterns, and practical impact remain poorly understood. Through a mixed-methods approach analyzing 7,310 rules from 83 open-source projects alongside survey data from 99 developers, this work proposes the first systematic classification framework comprising five top-level categories and 25 subcategories. It reveals a significant discrepancy between developer intent and actual rule configurations, demonstrates that rule evolution is frequent and primarily driven by contextual expansion, and quantifies a substantial improvement in software compliance—increasing from an average of 49.14% to 72.13% (+22.99%)—following rule updates.