code-enabled time series analysis

Designs and implements reproducible, code-driven analyses and pipelines for temporal (time-series) data that combine raw data with executable code or code-executing agents (including LLM coding agents) to perform trend detection and estimation, periodicity analysis, statistical tests, rolling-window computations, and other multi-step analytic workflows. Builds iterative query-and-execute processes that run computations, produce summaries and visualizations, and validate empirical time-series findings.

code-enabledtimeseriesanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study systematically evaluates the temporal reasoning capabilities of large language models (LLMs) in domains such as finance, healthcare, and environmental monitoring. It presents the first comparative analysis between two paradigms: direct processing of raw numerical time series and using LLMs as coding agents that generate executable Python code. A strong LLM-based adjudicator is introduced to conduct in-depth analysis of reasoning strategies and failure modes. Leveraging two established benchmarks and combining statistical testing with qualitative assessment, the experiments reveal that the coding-agent approach improves accuracy by up to 10% over pure textual input for comprehension tasks, yet overall error rates remain substantial at 22–34%. While LLMs often select appropriate statistical methods, they frequently overlook critical details. This work establishes the first systematic empirical framework and provides key insights into deploying LLMs for time series analysis.

automated decision-makingcoding agentslarge language models

Existing large language model (LLM)-driven data analysis tools are often confined to isolated subtasks and struggle to support end-to-end executable analytical workflows. This work proposes an autonomous, sandboxed, and auditable end-to-end system that leverages LLMs for action planning, iteratively generating structured operations, executing code in a secure environment, and integrating streaming traceability with intermediate result previews. By unifying a structured action backend, sandboxed execution, and an interactive visual interface—features integrated here for the first time—the system enables users to drive complete analytical workflows using only natural language. Users can inspect, modify, and export the entire process and its outputs directly within a web browser, ensuring full reproducibility, editability, and transparency throughout the analytical pipeline.

action tracedata analysisend-to-end workflow

AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents

Dec 26, 2025
BP
Bhanu Prakash Vangala
🏛️ University of Missouri | SRI International

This study investigates the executability of code generated by large language models (LLMs) in clean, minimal environments, revealing a substantial gap between declared dependencies and actual runtime requirements. We propose a novel three-layer dependency framework—comprising declared, available, and runtime dependencies—to systematically quantify dependency inconsistencies in LLM-based programming agents and assess cross-language reproducibility. Using standardized prompt sets across Python, JavaScript, and Java, we conduct automated dependency parsing and environment validation on Claude Code, Codex, and Gemini. Results show that only 68.3% of generated projects execute out-of-the-box; execution success rates are 89.2% for Python and 44.0% for Java, highlighting language-specific disparities. On average, dependency graphs inflate 13.5× relative to declared dependencies, exposing pervasive implicit dependency issues. This work provides critical empirical evidence and a methodological foundation for improving the reliability and engineering deployability of LLM-generated code.

Investigates reproducibility of LLM-generated code in clean environmentsQuantifies dependency gaps using a three-layer framework across languagesReveals hidden dependencies causing execution failures in generated projects

Flowco: Rethinking Data Analysis in the Age of LLMs

Apr 18, 2025
SN
Stephen N. Freund
🏛️ Williams College | University of California, Los Angeles | University of Massachusetts Amherst | Amazon Web Services

In data science practice, while large language models (LLMs) can automatically generate analytical code, they lack support for fine-grained control, intermediate result validation, and iterative optimization—compromising analysis controllability, verifiability, and reproducibility. To address this, we propose Flowco: a hybrid, proactive visual dataflow programming framework that uniquely embeds LLMs throughout the entire analytical workflow—spanning code generation, debugging, validation, and iteration. Flowco integrates visual dataflow modeling, LLM-augmented collaborative reasoning, and traceable execution graphs. A user study demonstrates that Flowco significantly improves novices’ efficiency in constructing, debugging, and optimizing analytical tasks, while preserving usability and simultaneously ensuring controllability, verifiability, and reproducibility. By unifying human-in-the-loop interaction with LLM intelligence in a structured, auditable environment, Flowco establishes a novel paradigm for democratizing data science in the LLM era.

Enabling non-experts to conduct data analyses using LLMsProviding fine-grained control and verification in analysis stepsSupporting iterative refinement of data analysis workflows

ML scientists struggle to efficiently audit the iterative coding processes of LLM-based programming agents; existing manual inspection of individual outputs hinders tracing code evolution, comparing multi-turn behaviors, and identifying improvement opportunities. Method: We propose the first multi-level visual analytics system specifically designed for LLM programming agents, enabling coordinated analysis across three granularities: code (diff-based comparison), process (reconstruction of solution paths), and model (cross-model behavioral contrast). Built upon the AIDE framework, the system integrates interactive visualization techniques to ensure traceable coding trajectories, visible iterative differences, and comparable model strategies. Contribution/Results: Evaluated on multiple Kaggle competition case studies, our system significantly enhances users’ depth of understanding of agent behavior and accelerates prompt debugging. It establishes a novel paradigm for controllable development and explainability research of LLM agents.

Enabling comparative analysis across code, process, and LLM levelsEnhancing understanding of LLM coding agent behaviorsImproving efficiency in reviewing iterative code generation

Latest Papers

What's happening recently
View more

This work addresses the challenges of low reliability, poor auditability, high cost, and security risks associated with runtime invocation of large language models (LLMs) in high-stakes enterprise workflows. To overcome these issues, we propose a novel “compiled AI” paradigm, wherein LLMs generate executable code during compilation, eliminating the need for model calls at runtime and thereby ensuring deterministic execution. We present the first systematic application of this paradigm to high-risk scenarios, integrating constrained code generation, a four-stage verification pipeline, template-embedded business logic functions, and an operation-oriented evaluation framework to jointly achieve reliability, auditability, and security. Experiments demonstrate a 96% success rate on function-calling tasks with zero runtime token consumption; 80.0% and 80.4% accuracy on critical field extraction and line-item recognition in document intelligence tasks; and strong security performance, with 96.7% prompt injection detection accuracy and 87.5% static analysis precision without false positives.

auditabilitycompiled AIdeterministic execution

Existing coding agents are largely confined to code generation and lack support for the full workflow lifecycle, including composition, iteration, deployment, and sharing. This work proposes CURATE, a novel system that integrates modular cataloging and FAIR principles into a large language model–driven multi-agent framework to enable human-in-the-loop, end-to-end workflow development and automated execution. Built upon Claude Opus 4.8, CURATE incorporates user-in-the-loop mechanisms and a module registry to facilitate cross-workflow sharing of reusable components. The system successfully reproduces four SeBS-Flow benchmark workflows and automatically constructs a complex anaerobic digestion simulation pipeline, demonstrating its feasibility and effectiveness in supporting comprehensive, collaborative scientific workflow automation.

code generationdeploymentmodule reuse

This work addresses the lack of realistic coding agent workload data in current large language model (LLM) serving systems, which hinders efficient optimization. We present the first large-scale trace dataset from real-world coding agents, comprising approximately 4,300 sessions, 350,000 LLM inference steps, and 430,000 tool invocations. The dataset reveals key runtime characteristics in multi-agent, multi-model scenarios, including long autonomous loops, short outputs paired with long contexts, a heavy-tailed tool invocation distribution, and high—but imperfect—prefix cache hit rates. Leveraging these insights, we propose targeted optimizations: low-overhead tool invocation, context-aware prefilling, semantic tool latency prediction, and improved KV cache management. Our findings provide empirical grounding and new directions for designing efficient LLM serving systems.

coding agentsLLM servingserving optimization

This study addresses the challenge that existing AI code generation tools often fail to ensure fidelity in the software implementation of statistical methods, thereby introducing implementation distortions. To mitigate this issue, the authors propose a novel multi-agent development paradigm built upon Claude Code, incorporating an information isolation mechanism. In this framework, a planning agent generates separate specifications for implementation, simulation, and testing, which are then executed by dedicated agents operating in mutual isolation. This approach pioneers the use of information barriers in AI-assisted programming, eliminating reliance on prior knowledge in code generation while preserving researchers’ full control over methodological decisions. Empirical evaluations demonstrate that the workflow successfully implements probit estimation and integrates with multiple R and Python statistical packages, effectively offloading engineering overhead without compromising implementation accuracy.

AI code generationfaithful implementationquantitative research

Hot Scholars

HC

Haeran Cho

University of Bristol
change-point detectionnonstationary time series analysishigh-dimensional data analysisenergy data modelling
TC

Tanujit Chakraborty

Associate Professor of Statistics and Data Science at Sorbonne University
Machine LearningTime Series ForecastingSpatial StatisticsHealth Data Science
YL

Yuxuan Liang

Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Spatio-Temporal Data MiningUrban ComputingUrban AIFoundation Models
NT

Neil Thompson

Director, MIT FutureTech at Computer Science and A.I. Lab and the Initiative on the Digital Economy
Moore's Law and Computer PerformanceTools and InnovationPatenting & LicensingExecuting on