Score
Designs and produces structured research roadmaps and agendas—documents that map current knowledge, identify unresolved gaps and cross-cutting challenges, prioritize open problems, and set future directions. Builds prioritized, actionable plans by analyzing literature and stakeholder inputs to recommend research priorities, co-design and integration strategies, evaluation protocols and benchmarks, and experimental and reproducibility practices.
This study addresses the limitations of large language models in generating scientific research roadmaps—specifically, insufficient domain expertise, suboptimal task decomposition, and logical inconsistencies—by proposing RoadMapper, a multi-agent collaborative framework that structures the generation process into three phases: initial drafting, knowledge augmentation, and iterative critique-revision-evaluation. The work introduces RoadMap, the first benchmark dataset for research roadmap generation, and integrates knowledge enhancement with multi-agent coordination. Experimental results demonstrate that RoadMapper significantly outperforms baseline methods in domain specificity, logical coherence, and practical utility, achieving an average performance improvement of over 8% while reducing generation time to merely 16% of that required by human experts.
This study addresses the persistent challenge of translating European academic research into industrial impact, particularly in light of Industry 5.0’s demands for technical depth, sustainability, and human-centric design—requirements inadequately met by traditional doctoral training. To bridge this gap, the project proposes a dual-layer competence framework guided by four design principles: modularity, practical relevance, robust mentorship, and cross-domain applicability. Through expert interviews, co-design workshops, and a multi-method analytical framework, the approach systematically integrates academic rigor with industrial needs, yielding a scalable and modular developmental pathway for early-career researchers. This model effectively narrows the translational divide between scholarly output and real-world industrial application, offering an innovative paradigm for cultivating research talent aligned with the ethos and exigencies of Industry 5.0.
Scientific software frequently suffers from poor robustness, low maintainability, and weak sustainability. To address these challenges, this work systematically integrates software engineering best practices with domain-specific research requirements, proposing a set of ten high-quality principles for building scientific software across its entire lifecycle. The principles cover critical phases—including project planning, readable coding, version control, automated testing, modular design, reproducibility assurance, performance optimization, and long-term maintenance—and are supported by technical enablers such as automated documentation generation, continuous integration, and performance profiling. Designed to be both broadly applicable and practically actionable, the framework has been empirically validated across multiple scientific domains. Results demonstrate significant improvements in software reliability, reusability, and collaborative efficiency within research communities, thereby enhancing the academic impact of scientific tools and advancing open science and reproducible research ecosystems.
Existing research ideation tools emphasize breadth-oriented idea generation but lack support for iterative refinement, elaboration, and evaluation—hindering literature-grounded, deep-reading–driven conceptual evolution. Method: We propose the first literature-driven interactive research ideation system, integrating a composable “idea element” canvas model with a multi-dimensional (problem/solution/evaluation/contribution) co-evolution mechanism. Our approach innovatively incorporates LLM-powered literature-aware feedback generation, graph-structured idea modeling, and interactive multi-path variant exploration. Contribution/Results: Experiments demonstrate a 42% increase in user-generated idea output and significantly enhanced detail elaboration. Seven researchers successfully applied the system across the full ideation pipeline—from initial topic conception to paper outline revision—validating its efficacy in supporting deep, iterative, literature-informed research design.
Scientific software development suffers from poorly specified requirements and inadequate management, severely compromising software quality and experimental reproducibility. To address this gap, this study formally establishes scientific software as a novel application domain for requirements engineering (RE). Through eight in-depth interviews with 12 researchers, we conduct an exploratory qualitative study employing thematic coding analysis. Our findings identify three core challenges: (1) highly ambiguous and evolving requirements, (2) latent or unidentified stakeholders, and (3) absence of systematic requirement validation mechanisms. Based on these insights, we propose a domain-specific RE vision and a challenge framework tailored to scientific software contexts. This work lays the theoretical foundation and methodological guidance for lightweight, agile, and traceable RE practices in scientific software development—thereby filling a critical void in systematic RE research for this domain.
This study addresses the limited depth and breadth of interdisciplinary knowledge integration caused by the poor reusability of underlying data in traditional literature reviews. We propose a "dual-track integration" conceptual framework that leverages the TIB Knowledge Loom to generate machine-readable outputs, combining systematic review methodologies with knowledge gap mapping techniques to comparatively evaluate manual extraction against automated approaches for knowledge synthesis. Our analysis reveals that only 8% of the examined literature provides reusable data, while demonstrating that constructing knowledge integration frameworks linking publications to research infrastructure effectively expands integration pathways. The core contribution of this work lies in identifying that ensuring the accessibility and executability of data, code, and workflows is essential for overcoming existing bottlenecks in knowledge integration.
This work addresses the limitation of existing research agents that oversimplify complex scientific projects into single tasks, resulting in ambiguous task boundaries, disorganized execution, and missing deliverables—challenges that hinder long-horizon, multi-objective, and dependency-sensitive research planning. To overcome this, the authors propose a graph-guided, project-level planning approach that explicitly decomposes a research project into executable task compositions with clearly attributed contributions and explicit dependencies, leveraging an innovative atomic representation and a directed provenance graph. A lightweight Bernoulli block model optimizes task selection, generating standardized task contracts that specify objectives, dependencies, and constraints, enabling seamless decoupled integration with arbitrary executors. Evaluated on ten scientific benchmarks, the method achieves an average quality score of 7.15, significantly outperforming baselines (4.58 and 5.31), and when integrated with AutoResearchClaw, boosts downstream task accuracy from 0.536 to 0.759.
This study addresses the fragmentation of evaluation criteria for automated research systems and the difficulty of direct cross-task comparison. Employing a systematic literature review, it comprehensively examines evaluation designs across six task categories, including literature synthesis and ideation. By comparing benchmark construction and scoring protocols, this work proposes a complementary evaluation framework encompassing output-level, process-level, and human-subject assessments. It reveals the capability differences reflected by distinct designs and underscores the critical role of calibration specificity and resource budgets in performance interpretation. Furthermore, the project identifies gaps in diagnostic evaluation and provides recommendations for standardized reporting and auditing. Ultimately, these contributions offer practical guidance for benchmark selection and future research design in evaluating automated scientific discovery systems.
This study addresses the challenges of disconnected planning and investigation, difficult evidence integration, and inefficient local revision in deep research report generation. To this end, it proposes a multi-agent collaborative framework built upon an enhanced LongCat model. Methodologically, the work introduces a structured ResearchSpec to decouple global planning from parallel chapter-level investigation, and designs a globally guided, chapter-level coordinated revision mechanism that avoids full-document rewriting while supporting automated training data construction. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on benchmarks such as DeepResearchBench II, significantly improving both the readability and generation efficiency of the produced reports.
This work addresses the lack of systematic grounding in existing research idea generation methods, which often fail to identify bottlenecks, differentiate prior work, or assess risks effectively. To bridge this gap, the authors propose an evidence-driven framework for generating research ideas, introducing “idea cards” that encapsulate 15 reusable creative patterns distilled from top-tier machine learning conference papers. Each card is structured around context, bottleneck type, and differentiation strategy. The framework integrates multi-source literature retrieval, prior-work collision detection, and pattern-guided generation to support traceable and auditable proposal development. In blind evaluations, proposals generated by this approach significantly outperformed both unskilled and general-purpose baselines in quality while maintaining high novelty.