Score
Synthesizing current limitations, open challenges, and actionable future directions to prioritize research agendas and deployment requirements (datasets, models, trustworthy technologies) for a given technical area.
Scientific software development suffers from poorly specified requirements and inadequate management, severely compromising software quality and experimental reproducibility. To address this gap, this study formally establishes scientific software as a novel application domain for requirements engineering (RE). Through eight in-depth interviews with 12 researchers, we conduct an exploratory qualitative study employing thematic coding analysis. Our findings identify three core challenges: (1) highly ambiguous and evolving requirements, (2) latent or unidentified stakeholders, and (3) absence of systematic requirement validation mechanisms. Based on these insights, we propose a domain-specific RE vision and a challenge framework tailored to scientific software contexts. This work lays the theoretical foundation and methodological guidance for lightweight, agile, and traceable RE practices in scientific software development—thereby filling a critical void in systematic RE research for this domain.
Centralized biological data repositories face single-point failure risks—including cyberattacks, natural disasters, and governance or funding disruptions—jeopardizing data availability, integrity, and research continuity. To address this, we propose a hybrid scientific data infrastructure integrating federated architecture with decentralized technologies. Our approach employs distributed storage, cross-domain federated governance, and on-chain/off-chain协同 mechanisms for data integrity verification. The resulting framework enhances resilience and governance equity while adhering to FAIR principles. It significantly reduces dependence on central authorities, promotes fairer global data sovereignty distribution, and improves long-term sustainability. Empirical evaluation demonstrates robust fault tolerance, scalable interoperability across heterogeneous domains, and verifiable provenance tracking. This infrastructure provides a resilient foundation for open science, enabling trustworthy, persistent, and collaboratively governed biological data stewardship.
This study addresses persistent challenges in scientific communication and aerospace engineering—namely data silos, insufficient collaboration incentives, and legal barriers—that hinder the implementation of FAIR (Findable, Accessible, Interoperable, Reusable) principles. To overcome these limitations, this work proposes a novel, scalable knowledge infrastructure framework that integrates human–AI collaboration, knowledge graphs, and user-centered design across technological, social, and legal dimensions. The framework encompasses automated information processing workflows, a wiki-style digital library, and demand-driven interactive interfaces. Pilot implementations demonstrate its effectiveness in consolidating fragmented knowledge resources and establishing a viable collaborative paradigm for sparsely networked domains. Nevertheless, institutional and sociocultural barriers remain significant and require further intervention to fully realize the framework’s potential.
This study addresses the fragmented landscape of open access (OA) monitoring tools—characterized by inconsistent metrics, heterogeneous data standards, and limited cross-platform comparability. To tackle this, we conducted a comprehensive survey of nearly 60 international OA dashboards, designed the first domain-specific metadata schema for OA monitoring, and constructed a structured, extensible dataset. Methodologically, we introduced a participatory curation mechanism and a systematic indexing framework to enable sustainable, multi-stakeholder collaboration in data maintenance. The outcome is the first global, comprehensive, and open-source OA dashboard dataset. It provides empirical foundations and standardized analytical infrastructure for research management bodies, policymakers, and library and information science scholars. By harmonizing indicators and ensuring transparency, scalability, and reproducibility, this work significantly enhances the comparability, transparency, and long-term sustainability of OA progress monitoring. (149 words)
Rapid AI advancement poses novel governance challenges, necessitating a rigorous, technically grounded approach to AI governance. Method: This work introduces “technical AI governance” as a distinct paradigm and establishes the first interdisciplinary analytical framework—integrating AI safety, mechanism design, policy modeling, and governance theory—to systematically address three core problem domains: risk identification, evaluation of intervention effectiveness, and compliance mechanism design. Adopting a problem-driven methodology, it clarifies how technical tools can concretely support governance practice. Contributions/Results: (1) A formal, structured definition of technical AI governance and a taxonomy of its core problems; (2) The first publicly available, extensible open-problems catalog for technical AI governance, bridging methodological gaps between technical and policy communities; and (3) An actionable, problem-oriented investment guide for researchers and funding agencies to prioritize high-impact technical governance research.
This work proposes an automated method for constructing large-scale technology roadmaps to uncover dependency and evolutionary relationships among scientific contributions. Leveraging advanced natural language processing techniques, the authors extract 2 million scientific contributions from 230,000 open-access papers and build a structured dependency graph comprising 12.5 million prerequisite edges. They further introduce, for the first time, the task of "scientific prerequisite prediction" and demonstrate the feasibility of their approach by achieving a mean average precision (MAP) of 0.48 on this task. This study delivers the first large-scale dependency graph of scientific contributions, establishing a novel paradigm for assessing scientific impact and enabling automated knowledge discovery.
Empirical research on scholarly software lacks large-scale, evidence-based foundations. Method: We constructed the largest literature-linked open-source research software dataset to date, comprising 134,352 distinct projects and 134,154 source code repositories, along with their citations in open-access publications. By systematically integrating metadata from open publishing platforms and code hosting services, we extracted structured information—including version history, licenses, programming languages, and functional descriptions—enabling the first fine-grained mapping between research software and its associated scholarly outputs. Contribution/Results: The publicly released dataset includes complete metadata for over 120,000 projects, substantially addressing the scarcity of high-quality empirical data in research software engineering (RSE). It provides a reproducible foundation for assessing software impact, analyzing development practices, and informing evidence-based policy formulation in scholarly software infrastructure.
Scientific data often require extensive manual curation before being usable for scientific AI, lacking a unified framework for automated conversion, readiness assessment, provenance tracking, and agent integration. This work proposes REDI, an open-source framework that automatically transforms raw scientific data into AI-ready formats through a five-stage, fully traceable pipeline—ingestion, preprocessing, transformation, structuring, and output—while exposing the resulting workflows as callable skills for AI agents. REDI is the first framework to unify these capabilities; its companion tool, SetGo, ensures FAIR compliance and enables automatic catalog publishing. Leveraging parallel distributed processing and I/O performance profiling, REDI demonstrates effectiveness across climate science, proteomics, materials science, and nuclear fusion, with the climate use case achieving near-ideal strong scaling up to 100 nodes on the Frontier supercomputer.
This study addresses core bottlenecks in AI-powered scientific discovery—low data trustworthiness, poor model transferability, and the absence of an experimental-computational closed loop—by proposing a tripartite AI-driven research paradigm: “trustworthy data—transferable models—generative systems.” Methodologically, it integrates large foundation models with physics-informed generative modeling (e.g., electronic structure and synthetic feasibility constraints), active learning, and self-driving laboratory technologies to enable end-to-end, interpretable, and reproducible scientific workflows across biology, chemistry, materials science, climate science, and physics. Its key contribution is the first multi-disciplinary framework unifying AI, experimentation, and simulation, significantly enhancing physical consistency of models and autonomous scientific reasoning. This provides a systematic, transparent, efficient, and verifiable pathway toward AI-augmented scientific discovery.
This study addresses the significant challenge posed by the heterogeneity of space biology data, which severely limits the application of artificial intelligence (AI) in space life sciences. To overcome this barrier, the authors propose a three-tiered data restructuring framework—“FAIR → AI-ready → space-ready”—that systematically enhances data usability for AI through standardization, enriched metadata, purpose-built AI interfaces, and a distributed governance architecture. The work innovatively outlines an AI-ready data evolution pathway tailored for deep space exploration and advocates for the establishment of a neutral international coordinating body to ensure the trustworthiness, interoperability, and agent-accessibility of space biology data infrastructures. This approach provides both a technical roadmap and a governance blueprint to support multimodal AI applications in space biology.