Score
Designs and builds systems that translate natural-language questions into structured, executable analysis plans and interactive workflows, including NL-to-analysis planners and co-scientist agents. Implements dialog-state management, reasoning over intermediate results, and automated recommendations for follow-up analyses to support multi-turn analytic sessions.
Traditional data analysis methods face limitations in semantic understanding, natural language interaction, heterogeneous data processing, and autonomous workflow orchestration. Method: This paper systematically surveys recent advances in large language models (LLMs) and intelligent agent technologies for data analysis, proposing five design principles—semantic awareness, multimodal fusion, autonomous pipeline construction, tool-augmented workflows, and open-task support—to architect an intelligent analytical agent for cross-modal data (structured, semi-structured, unstructured, and heterogeneous). Core techniques include NL2GQL translation, chart and document understanding, cross-modal alignment, and scalable tool invocation mechanisms. Contribution/Results: The study clarifies the technical evolution and application paradigms of LLM-driven data analysis, identifies critical challenges—including interpretability, robustness, and domain adaptation—and outlines future research directions. It provides a systematic foundation for both theoretical advancement and engineering practice in intelligent data analytics systems.
Service workflows in customer service dialogues are often missing and unstructured, leading to inconsistent AI responses. Method: This paper proposes the first integrated framework for automatic dialogue workflow extraction and simulation-based evaluation. It combines retrieval-augmented generation (RAG) with a novel question-answering–based chain-of-thought (QA-CoT) prompting technique to improve structured workflow generation accuracy. Additionally, it introduces a scalable two-agent (agent + customer) simulation evaluation mechanism for automated, large-scale workflow assessment. Contribution/Results: We design a macro-accuracy metric that achieves high agreement with human evaluation (Spearman ρ > 0.92). On the ABCD and SynthABCD datasets, our method improves average macro-accuracy by 12.16%, significantly mitigating response inconsistency caused by workflow omission in service AI systems.
Entry-level data analysts struggle to perform advanced text analysis using professional NLP tools due to high technical barriers. Method: This paper proposes a human-in-the-loop text analytics system built upon a “decompose–execute–evaluate” closed-loop framework: (1) it introduces the first MCTS-based generative reasoning guidance with explicit human feedback integration; (2) it automatically synthesizes executable text analytics pipelines; and (3) it jointly leverages LLM-based automated evaluation and interactive visual validation. Contribution/Results: The system significantly lowers the NLP adoption barrier while enhancing interpretability and controllability of analysis. Quantitative experiments and a user study involving 24 participants—from novices to experts—demonstrate that non-expert users achieve >89% task success rates and exhibit a 42% improvement in error identification, markedly advancing usability, effectiveness, and trustworthiness of text analytics.
This work addresses the unreliability of large language models (LLMs) in executing structured workflows specified through natural language. To overcome this limitation, the authors propose RunAgent, a multi-agent platform that integrates the expressiveness of natural language with the determinism of programmatic execution through a novel agent language. RunAgent introduces a constraint-guided stepwise execution mechanism augmented with explicit control structures, dynamic selection of reasoning strategies, and context filtering to ensure robust task execution. The framework automatically derives verifiable constraints and supports a hybrid paradigm combining tool invocation, Python code generation, and execution. Evaluated on the NaturalPlan and SciBench benchmarks, RunAgent substantially outperforms both baseline LLMs and the current state-of-the-art PlanGEN method.
Qualitative data analysis in software engineering faces challenges including time intensity, poor reproducibility, and difficulty ensuring inter-rater reliability; the potential of large language models (LLMs) for human–AI collaboration in such tasks remains underexplored. This paper introduces the first explainable multi-agent framework tailored for qualitative research, enabling automated coding, theme extraction, and cross-textual synthesis via role-based task decomposition, prompt engineering, iterative validation, and a closed-loop human feedback mechanism. The architecture preserves human oversight and ensures analytical traceability, overcoming LLM limitations in low-shot, high-reliability settings. Empirical evaluation demonstrates a 3.2× improvement in analysis efficiency, scalability to hundreds of interviews, 89.7% accuracy in theme identification, and strong endorsement by domain experts.
Current LLM-based dialogue agents face three key challenges in knowledge-intensive, task-oriented dialogues: frequent hallucinations, weak conditional logic reasoning, and difficulty integrating heterogeneous knowledge sources—leading to low reliability. This paper introduces Genie, a framework accompanied by the declarative GenieWorksheet specification, which enables robust parsing of complex conditional instructions and consistent multi-source knowledge integration via controllable strategy programming, strong knowledge grounding, and collaborative LLM execution. Our approach unifies programmable architecture design with declarative policy modeling to significantly improve instruction-following stability. On the STARV2 benchmark, Genie outperforms prior state-of-the-art by 20.5%. In real-user experiments, it achieves 21.1%, 20.1%, and 61% improvements over GPT-4+Function Calling in execution accuracy, dialogue act accuracy, and goal completion rate, respectively—demonstrating the effectiveness and practicality of highly reliable task–knowledge fusion agents.
Quantitative syntactic research often suffers from poor auditability, limited shareability, and hindered iteration due to computational logic being embedded implicitly within scripts. This work proposes QLWF, a platform that leverages an AI-assisted five-stage pipeline to transform natural language research descriptions into explicit, executable workflows. Its key innovation lies in integrating visual logic restructuring with deterministic execution semantics, enabling incremental iteration driven by localized modifications without requiring full workflow regeneration. Built upon large language models, QLWF ensures reproducibility through a fixed node library and a dedicated execution engine, while an incremental refinement mechanism reduces restructuring overhead. Experimental results demonstrate that on the 64-task QL-Bench, the platform generates workflows that are 100% structurally valid and 98.4% semantically reasonable; in a 12-task lifecycle evaluation, incremental revisions achieved a 100% success rate with only one-third of the token consumption required for full regeneration.
This work addresses the lack of semantic automation in scientific workflow construction, which traditionally relies on experts to manually translate research questions into formal specifications. The authors propose a three-layer agent architecture: a large language model (LLM) parses natural language queries into structured intents; domain-specific “skill” documents encode terminology mappings and parameter constraints; and a validated generator produces reproducible DAG-based workflows. By confining LLM uncertainty to the intent extraction phase and integrating domain knowledge through the skill layer, the approach ensures workflow consistency while significantly improving semantic accuracy and execution efficiency. Experiments on 150 queries show that skills increase exact intent match accuracy from 44% to 83% and reduce data transfer by 92%. The end-to-end pipeline executes in under 15 seconds on average in Kubernetes, with a per-run cost below $0.001.
Traditional structured tool-calling approaches in large language model (LLM) agents suffer from low reliability, high error rates, and substantial computational overhead. This work proposes the Natural Language Tool (NLT) framework and, through 8,560 experiments spanning 14 models, presents the first independent reproduction and open-source validation of its efficacy. The study demonstrates that model capability significantly modulates NLT performance and quantifies its deployment-level advantages: an absolute improvement of 14.9 percentage points in overall tool-calling accuracy, a 93% reduction in critical errors, and a 25.2% decrease in token consumption. These gains are especially pronounced for models lacking native tool-calling support or those of smaller scale.