Research with AI Agents: How Agentic Systems Are Changing Scientific Work

πŸ“… 2026-09-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unclear conditions under which AI agents yield reliable efficiency gains in scientific research and the ambiguous boundaries of human responsibility. To investigate this, it systematically evaluates agent performance on standardized tasks such as literature retrieval and data analysis, proposing a component-level verification strategy to resolve the unauditability of complex reasoning chains. Furthermore, it constructs an end-to-end framework encompassing automated subtask decomposition, code execution, and clinical decision support. The findings demonstrate significant efficiency improvements in standardized tasks while revealing insufficient reliability in model interpretability. Consequently, this work establishes that researchers must retain ultimate accountability and affirms the centrality of expert judgment in assessing social value, thereby providing a theoretical foundation for optimizing human-AI collaboration paradigms.
πŸ“ Abstract
Background. Agentic AI systems independently decompose tasks such as literature search, data analysis, and programming into subtasks, search the web, access databases, and execute code. This allows them to perform digital research tasks at high speed. Objectives. Under what conditions does the use of agentic systems produce reliable efficiency gains, and which tasks remain with researchers? Materials and methods. Summary of current studies on literature searches, data analysis, software development, and clinical decision support. Results. For digital activities, work shifts from execution to steering and review. Efficiency gains are greatest when expected behavior can be formalized in advance and tested automatically. In complex agentic systems, recorded sequences of reasoning steps and tool calls can quickly become too extensive for human review. Furthermore, explanations generated by the model do not reliably reflect how an output was produced. One possible step toward more reliable systems is the validation of individual components. The limited reviewability extends beyond research itself; the review of scientific articles and grant proposals is also reaching capacity limits. In pathology, curated and annotated data, researchers' own analytical skills, and institutional exchange of experience are becoming increasingly important. Conclusions. Researchers remain responsible for their results. They must determine what to delegate and how to review the results. The importance of a research question to patients, the field, and society cannot be fully assessed using formalized criteria and remains a matter of expert judgment. Agentic systems can free up time for this.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Scientific Research
Efficiency Gains
Task Delegation
Human Oversight
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic AI
Task Decomposition
Automated Validation
Component Verification
Human-AI Collaboration
Johannes Lotz
Johannes Lotz
Fraunhofer MEVIS
Image RegistrationDigital Pathology
M
Markus Wenzel
Fraunhofer Institute for Digital Medicine MEVIS, LΓΌbeck, Germany; Constructor University, Bremen, Germany