multi-agent data integration

Designs and implements systems where multiple autonomous agents coordinate to extract, transform, reconcile, and merge heterogeneous datasets into a single integrated repository, specifying agent roles, communication and orchestration protocols, and pipeline stages. Builds and analyzes workflows for data curation, conflict resolution of overlapping signals, provenance capture during merges, and continual synchronization between sources.

multi-agentdataintegration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

本文提出数据代理新范式,通过自主执行数据任务解决传统数据系统在AI时代面临的限制,包括语义理解、自主编排和主动处理等方法。

autonomous orchestrationdata systemsheterogeneous data

This work proposes a hierarchical meta-agent architecture to overcome the static and rigid nature of traditional data processing pipelines, which struggle to autonomously monitor and optimize themselves post-deployment. The framework integrates three core components—planning, agent orchestration, and a monitoring feedback loop—to enable dynamic construction, execution, and iterative refinement of end-to-end workflows. It introduces context-aware optimization, adaptive workload partitioning, and progressive sampling mechanisms, facilitating agent reuse and seamless integration with external tools. These innovations significantly enhance system flexibility and scalability. Experimental results demonstrate the framework’s effectiveness and practicality in automatically constructing, continuously monitoring, and adaptively optimizing diverse data processing tasks.

adaptive pipeline optimizationautonomous data processingcontext-aware data processing

Autonomous Data Agents: A New Opportunity for Smart Data

Sep 23, 2025
YF
Yanjie Fu
🏛️ Arizona State University | University of Kansas | University of Notre Dame | Duke University

Current data processing remains heavily manual, suffering from labor intensity, poor scalability, and structural misalignment with AI-driven applications. To address these limitations, this paper proposes DataAgents—a paradigm for autonomous data agents powered by large language models (LLMs). DataAgents enables end-to-end automation from unstructured data to actionable knowledge through task decomposition, action reasoning, code generation, and coordinated invocation of heterogeneous tools. Its core innovations include dynamic workflow planning and cross-task adaptive mechanisms, supporting intelligent, unified processing across the full data lifecycle—including acquisition, integration, cleaning, transformation, and augmentation. Experimental evaluation demonstrates significant improvements in intelligence, generalization, and scalability of data operations. By bridging semantic and operational gaps in data processing, DataAgents advances the “data-to-knowledge” paradigm toward practical, robust, and scalable realization.

Managing growing data complexity through autonomous AI systemsOptimizing data workflows for enhanced AI utilizationTransforming unstructured data into actionable knowledge efficiently

Earth science data are growing rapidly, yet their reusability remains limited. To address this challenge, this work proposes PANGAEA-GPT—a hierarchical multi-agent system featuring a centralized supervisor-worker architecture. The system integrates data-type-aware routing, sandboxed deterministic code execution, and an execution-feedback-driven self-correction mechanism to autonomously orchestrate complex, multi-step analytical workflows. Evaluated in physical oceanography and ecology scenarios, PANGAEA-GPT enables end-to-end data analysis with minimal human intervention, substantially enhancing the discoverability and usability of heterogeneous geoscientific datasets.

autonomous discoverydata reusabilitygeoscientific data archives

A Survey on Agent Workflow -- Status and Future

Aug 02, 2025
CY
Chaojia Yu
🏛️ Sichuan University

In the era of large language models, agent workflows face critical challenges in scalability, controllability, and security. To address these, this paper presents a systematic literature review and proposes, for the first time, a dual-dimensional taxonomy—spanning functional capabilities (task planning, multi-agent collaboration, tool integration) and architectural characteristics (role definition, orchestration process, specification languages). Through comparative analysis of over twenty representative academic and industrial systems, we identify recurring design patterns and persistent technical bottlenecks. We further introduce security-enhanced orchestration optimization strategies and pinpoint core gaps, including the lack of standardization and insufficient multimodal integration. This work establishes a foundational theoretical framework and practical guidelines for the design, evaluation, and evolution of agent workflows, advancing the field toward structured, trustworthy, and multimodal-cooperative paradigms.

Addressing workflow optimization, security, and future challengesClassifying systems by functional and architectural featuresSurveying agent workflow systems for scalable AI behaviors

Latest Papers

What's happening recently
View more

This study addresses the high costs, limited cross-domain generalization, and reliance on manual rules in dataset curation by proposing a multi-agent collaborative framework. The framework decouples the curation pipeline into five programmable stages and introduces a pioneering parallelized workflow that integrates online knowledge retrieval with skill evolution modules. Through multi-agent orchestration, context exploration, dynamic metric computation, and self-evolution mechanisms, it enables automated evaluation and filtering. Experimental results demonstrate that the proposed method reduces data noise rates by 36.03 percentage points and improves downstream model F1 scores by 8.88 percentage points, significantly enhancing the flexibility and adaptability of the curation process.

AutomationData QualityDataset Curation

This study addresses the silent failures and insufficient reliability of LLM-based data agents caused by the absence of workflow constraints, proposing a workflow-centric, five-stage data agent framework. It first establishes a taxonomy of data agents to identify four critical reliability gaps. Subsequently, it introduces a shared verification-repair paradigm that integrates fifteen technical strategies—including structure probing and state reconstruction—alongside a workflow harness mechanism, thereby achieving closed-loop control from perception to remediation. Finally, this work presents a comprehensive evaluation benchmark and releases an open-source code repository. The proposed framework significantly enhances both the robustness and automation level of end-to-end data analysis.

Data AgentsLarge Language ModelsReliability

Existing agent benchmarks struggle to evaluate the capability of automating end-to-end data science workflows in realistic environments. This work proposes the first benchmark that assesses agents’ performance on complete data science tasks within a real operating system, encompassing the full lifecycle—including data preprocessing, modeling, and visualization—and supporting coordinated use of multiple tools such as notebooks, terminals, and web browsers. The benchmark introduces a deterministic verification mechanism based on both intermediate outputs and final results. Evaluation across 275 tasks with 15 state-of-the-art models reveals that even the strongest closed-source model achieves only a 56.7% success rate, while open-source models generally fall below 1%, highlighting significant limitations in current agents’ multi-step reasoning and tool orchestration capabilities.

agent benchmarkingdata-science workflowsend-to-end automation

This study addresses the fundamental mismatch between the probabilistic interactions of large language model agents and the policy-driven architecture of data spaces, which impedes their seamless integration. To overcome this challenge, we propose Eunomia, a mediation layer architecture built upon the Model Context Protocol (MCP). Without requiring modifications to existing components, this approach translates data space capabilities into schema-driven, structured tools, thereby enabling controlled interaction and standardized interoperability between AI agents and data spaces. A prototype implementation validates the end-to-end workflow from catalog discovery to service invocation. The results demonstrate that this architectural mediation effectively reconciles compliance, interoperability, and decoupling while strictly preserving governance constraints.

Data SpacesGovernanceIntegration

Hot Scholars

WS

Weiyan Shi

PhD Student, Singapore University of Technology and Design
Human-Centred AIMultimodal LLMs
EL

Eryun Liu

Zhejiang University
Computer VisionImage ProcessingBiometricsFingerprint
YW

Yong Wang

Assistant Professor, Nanyang Technological University
Data VisualizationHCIHuman-AI CollaborationFinTech
YL

Yunyao Li

Director of Machine Learning, Adobe Experience Platform
Natural Language ProcessingMachine LearningHuman Computer InteractionData Management
JM

John Murzaku

Applied Scientist, Adobe
Computational LinguisticsTheory of MindCognitive States