DataWeave: Deploying Human-LLM Analytics for Exploratory Structured Data Analysis

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of workflow inefficiency, schema mismatch, and semantic misinterpretation that data journalists encounter when exploring structured data. To overcome these issues, we propose a human-led collaborative analysis system that positions large language models (LLMs) as interactive partners rather than automated engines. By integrating conversational interaction, schema grounding, analytical planning, and executable query generation, the system enables users to inspect, correct, and steer LLM outputs throughout hypothesis evolution, thereby establishing design principles for trustworthy collaboration. A case study conducted on the IPEDS dataset validates the effectiveness of the proposed approach. This work offers critical insights into the deployment of trustworthy human-AI collaboration for structured data analysis.
📝 Abstract
Data journalism, the practice of using data analysis to surface newsworthy stories, depends increasingly on the ability of reporters and investigative journalists to uncover trends, disparities, and accountability narratives. In practice, exploring large structured datasets remains slow and brittle: journalists must navigate hundreds of variables across many datasets over years, understand data coding conventions, and write non-trivial analysis code while hypotheses evolve. Although LLMs are often touted as "ask in English, get SQL/answers," real newsroom workflows expose recurring failures, e.g., schema mismatches and drift, misread domain semantics and units, and silent assumptions. We present DataWeave, a system that addresses these needs by combining conversational interaction, schema grounding, analytical planning, and executable query generation to support exploratory analysis over structured data. Rather than treating LLMs as autonomous answer engines, DataWeave frames them as interactive partners whose outputs can be inspected, corrected, and steered as hypotheses shift. We present a case study with professional journalists using our system to analyze the U.S. Department of Education's Integrated Postsecondary Education Data System (IPEDS), a high-stakes public dataset with substantial domain semantics and frequent schema updates. We also report how deployment experience and iterative refinement shaped the current DataWeave architecture and its analytical workflow. Our findings distill design principles and deployment lessons for trustworthy human-LLM collaboration in structured data analysis.
Problem

Research questions and friction points this paper is trying to address.

exploratory data analysis
structured data
data journalism
large language models
human-LLM collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-LLM Collaboration
Schema Grounding
Exploratory Data Analysis
Data Journalism
Executable Query Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Raquib Bin Yousuf
Virginia Tech, VA, USA
H
Harith Laxman
Virginia Tech, VA, USA
V
Vitaliy Shkremetko
Virginia Tech, VA, USA
E
Eunice Son
Virginia Tech, VA, USA
S
Shambhavi Verma
Virginia Tech, VA, USA
B
Brian O'Leary
The Chronicle of Higher Education, Washington, DC, USA
V
Venketesh Subramony
The Chronicle of Higher Education, Washington, DC, USA
S
Sylvain Nazef
The Chronicle of Higher Education, Washington, DC, USA
J
Jacquelyn Elias
The Chronicle of Higher Education, Washington, DC, USA
R
Ron Coddington
The Chronicle of Higher Education, Washington, DC, USA
C
Chris Contakes
The Chronicle of Higher Education, Washington, DC, USA
Michael Riley
Michael Riley
Google, Inc
speech and natural language processing
Naren Ramakrishnan
Naren Ramakrishnan
Thomas L. Phillips Professor, Virginia Tech
ForecastingMachine LearningComputational epidemiologyRecommender systemsVisual analytics