From Research Question to Scientific Workflow: Leveraging Agentic AI for Science Automation

📅 2026-04-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of semantic automation in scientific workflow construction, which traditionally relies on experts to manually translate research questions into formal specifications. The authors propose a three-layer agent architecture: a large language model (LLM) parses natural language queries into structured intents; domain-specific “skill” documents encode terminology mappings and parameter constraints; and a validated generator produces reproducible DAG-based workflows. By confining LLM uncertainty to the intent extraction phase and integrating domain knowledge through the skill layer, the approach ensures workflow consistency while significantly improving semantic accuracy and execution efficiency. Experiments on 150 queries show that skills increase exact intent match accuracy from 44% to 83% and reduce data transfer by 92%. The end-to-end pipeline executes in under 15 seconds on average in Kubernetes, with a per-run cost below $0.001.

Technology Category

Planning, Routing, and Scheduling: Planning with Language ModelsCognitive Modeling & Cognitive Systems: Agent ArchitecturesNatural Language Processing: Code Generation / Program Synthesis from Natural Language

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic searchEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd work
📝 Abstract
Scientific workflow systems automate execution -- scheduling, fault tolerance, resource management -- but not the semantic translation that precedes it. Scientists still manually convert research questions into workflow specifications, a task requiring both domain knowledge and infrastructure expertise. We propose an agentic architecture that closes this gap through three layers: an LLM interprets natural language into structured intents (semantic layer); validated generators produce reproducible workflow DAGs (deterministic layer); and domain experts author ``Skills'': markdown documents encoding vocabulary mappings, parameter constraints, and optimization strategies (knowledge layer). This decomposition confines LLM non-determinism to intent extraction: identical intents always yield identical workflows. We implement and evaluate the architecture on the 1000 Genomes population genetics workflow and Hyperflow WMS running on Kubernetes. In an ablation study on 150 queries, Skills raise full-match intent accuracy from 44% to 83%; skill-driven deferred workflow generation reduces data transfer by 92\%; and the end-to-end pipeline completes queries on Kubernetes with LLM overhead below 15 seconds and cost under $0.001 per query.
Problem

Research questions and friction points this paper is trying to address.

scientific workflow
research question translation
semantic automation
workflow specification
science automation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic AI
Scientific Workflow Automation
LLM-based Intent Extraction
Skill-based Knowledge Layer
Reproducible DAG Generation