🤖 AI Summary
This study addresses the challenges of error-prone planning, execution failures, and intent deviation when large language models are applied to multimodal data analysis. To overcome these limitations, this work proposes a planning system equipped with self-criticism and dynamic evolution capabilities. The proposed approach is grounded in a metadata knowledge graph and incorporates typed logical plans for rigorous validation. Furthermore, a self-criticism algorithm is introduced to enable the reuse of error diagnostics and the accumulation of planning experience. Experimental evaluations conducted on two public multimodal datasets demonstrate that the proposed system significantly enhances both the accuracy and robustness of analytical tasks.
📝 Abstract
Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. However, LLM-generated plans are error-prone: a plan may silently compute something other than what was asked, fail during execution, or return a result that misses the question. This paper presents WeaveData, a multimodal data analysis system with self-critiquing and self-evolving LLM plans. First, WeaveData generates a typed logical plan for each question and critiques it step by step before execution, and it checks the executed result against the question afterwards. Second, WeaveData evolves a plan that fails or misses the question: it diagnoses the failure with the actual data, reuses the results that remain valid, and accumulates planning experience for later questions. Third, WeaveData grounds planning in a metadata knowledge graph of all modalities, clarifies ambiguous questions with the user, and backs every model judgment with evidence in an interactive notebook. We demonstrate WeaveData on two public multimodal datasets.