Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the silent failures and insufficient reliability of LLM-based data agents caused by the absence of workflow constraints, proposing a workflow-centric, five-stage data agent framework. It first establishes a taxonomy of data agents to identify four critical reliability gaps. Subsequently, it introduces a shared verification-repair paradigm that integrates fifteen technical strategies—including structure probing and state reconstruction—alongside a workflow harness mechanism, thereby achieving closed-loop control from perception to remediation. Finally, this work presents a comprehensive evaluation benchmark and releases an open-source code repository. The proposed framework significantly enhances both the robustness and automation level of end-to-end data analysis.
📝 Abstract
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at https://github.com/DEEP-PolyU/Awesome-Data-Agents.
Problem

Research questions and friction points this paper is trying to address.

Data Agents
Reliability
Workflow Harnesses
Silent Failures
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Agents
Workflow Harnesses
Reliability
Silent Failures
Taxonomy
🔎 Similar Papers
No similar papers found.