design prompt chains

Design and build multi-stage prompt pipelines that decompose a task into ordered processing stages, creating stage-specific prompt templates and conventions for passing context and state between stages; include mechanisms to validate, correct, and recover from incorrect intermediate outputs to ensure end-to-end reliability.

designpromptchains

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Making Prompts First-Class Citizens for Adaptive LLM Pipelines

Aug 06, 2025
UC
Ugur Cetintemel
🏛️ Brown University

Current LLM pipelines treat prompts as unstructured strings, suffering from opacity, poor reusability, and lack of runtime controllability. This paper introduces SPEAR—the first framework to elevate prompts to structured, programmable, and versioned first-class citizens. Its core innovations are (1) a prompt algebra enabling composability, abstraction, and optimization; and (2) a versioned view mechanism supporting dynamic adaptation and fine-grained observability. SPEAR integrates operator fusion, prefix caching, and view reuse, and supports three optimization modes: manual, assisted, and fully automatic. Experiments demonstrate that SPEAR significantly outperforms static prompting and retry-based baselines, improving response quality by +12.7% in BLEU/accuracy and reducing inference latency by 38%. This work is the first to empirically validate the feasibility and effectiveness of structured, prompt-level optimization.

Enabling runtime prompt refinement based on execution signalsMaking prompts structured and adaptive in LLM pipelinesSupporting structured prompt management with versioning and logging

Prompt2DAG: A Modular Methodology for LLM-Based Data Enrichment Pipeline Generation

Sep 16, 2025
AA
Abubakari Alidu
🏛️ University of Milan-Bicocca

To address the high barrier to entry and heavy reliance on engineering expertise in data pipeline development, this paper proposes a hybrid generative approach that automatically compiles natural language specifications into reliable, executable Apache Airflow DAGs. The method synergistically integrates large language models (LLMs) for semantic understanding, structured template engines for deterministic syntactic constraints, and a multi-stage validation mechanism—balancing expressiveness with correctness guarantees. We introduce a novel three-dimensional evaluation framework—SAT (Semantic Accuracy), DST (Structural Integrity), and PCT (Programmatic Executability)—to systematically quantify generation quality. Experimental results demonstrate a 78.5% successful generation rate, substantially outperforming pure-LLM (66.2%) and end-to-end generative baselines (29.2%), while achieving over twofold improvement in cost efficiency per generated DAG. This work provides a practical, production-ready pathway toward democratizing low-code data pipeline development.

Automating data enrichment pipeline generation from natural languageBalancing flexibility and reliability in automated DAG creationEvaluating optimal LLM strategies for reliable workflow automation

This work addresses the high cost and sensitivity to phrasing inherent in manually crafted prompts, as well as the limited ability of existing automated optimization methods to systematically identify and correct failure patterns. To overcome these challenges, the authors propose Reflective Prompt Tuning (RPT), a novel framework that introduces a reflection mechanism leveraging large language model function calling. RPT employs a diagnostic function to analyze failure modes on an optimization set, generates structured reports, and iteratively refines prompts by integrating historical memory with confidence calibration. Experimental results demonstrate that RPT achieves performance gains of up to 12.9 points across three reasoning tasks, with particularly pronounced improvements in multi-hop and mathematical reasoning, while also enhancing the calibration of model output confidence.

automated promptingfailure patternsinstruction sensitivity

Tasks People Prompt: A Taxonomy of LLM Downstream Tasks in Software Verification and Falsification Approaches

Apr 14, 2024
VB
V. Braberman
🏛️ Universidad de Buenos Aires | CONICET-Universidad de Buenos Aires | The University of Manchester | Universidade Federal do Amazonas

Current LLM-native software engineering lacks a systematic practical framework—particularly in verification and falsification—necessitating unified task taxonomies and prompt-pattern conceptualizations. Method: We conduct a systematic literature review of over 100 papers, employing bibliometric analysis and conceptual clustering to map, classify, and abstract LLM-based downstream tasks in software engineering (SE). Contribution/Results: We propose the first fine-grained SE-specific taxonomy for LLM downstream tasks, encompassing six core clusters: testing, fuzzing, bug localization, vulnerability detection, static analysis, and program verification. Our taxonomy uniquely balances cross-task abstraction with task-specific variation modeling, uncovering generalizable prompt-engineering principles. It provides a foundational framework for targeted LLM adaptation, benchmark construction, and empirically grounded engineering practice in SE.

Developing conceptual frameworks for LLM-native software engineering practicesIdentifying compositional patterns for reliable LLM-native system designSystematically analyzing generative transformations in software verification

Latest Papers

What's happening recently
View more

This study addresses the frequent inefficiencies in human-AI collaboration caused by incomplete contextual information, which often leads to excessive iteration and suboptimal output quality. To mitigate this, the authors propose a structured context construction framework that integrates a five-role context package—comprising authority, exemplars, constraints, evaluation criteria, and metadata—within a four-stage workflow encompassing review, design, construction, and audit. Notably, this work pioneers the incorporation of information theory and reliability engineering principles into context quality assessment, yielding a reusable and auditable collaboration framework. Empirical results from 200 interaction trials demonstrate that the approach reduces the average number of iterations from 3.8 to 2.0, increases first-pass success rates from 32% to 55%, and achieves a final task success rate of 91.5%.

Context CompletenessHuman-AI CollaborationIteration Cycles

Automatically constructing high-quality, reusable skills from heterogeneous, fragmented interaction traces—often missing critical security behaviors—is highly challenging. This work proposes the W2S framework, which introduces a novel intermediate representation called RWSA to decouple skills into workflow structure, execution semantics, and runtime attachments, thereby enabling task decomposition, control-flow modeling, verification, rollback, and state management. W2S achieves efficient skill construction through trajectory segmentation, local skill draft generation, structural alignment, branch fusion, redundancy compression, and confidence-aware retention. Experimental evaluation across 70 skills demonstrates that W2S improves behavioral replay consistency by 10.5% compared to baseline approaches based on summarization and prompting.

agent trajectoriesinteraction tracesprocedural knowledge

Current prompt graphs lack a clear definition and standardized terminology, resulting in conceptual ambiguity in practice. This work addresses this gap by proposing a formal definition of prompt graph engineering through conceptual analysis, gray literature review, and systematic categorization. It identifies prompt graphs as first-class, executable, and improvable engineering artifacts and establishes four necessary constitutive conditions along with inclusion and exclusion criteria for operational validation. The proposed definition demonstrates consistent applicability across six major frameworks—including LangGraph and DSPy—thereby offering the field its first operational framework and shared vocabulary. Building on this foundation, the paper outlines a future research agenda structured around four key design tensions inherent to prompt graph development.

executable graphorchestration artifactprompt engineering

This study addresses the current lack of systematic research on operational frameworks and process mechanisms for AI software development agents. It proposes the first six-dimensional process taxonomy—encompassing specification, context, role, execution, validation, and portability—and employs targeted literature review, functional filtering, traction metrics, and a structured scoring rubric to conduct a multi-case comparative analysis of six representative frameworks. The analysis reveals a prevailing trend among mainstream frameworks toward de-emphasizing isolated prompts and instead reinforcing persistent artifacts and human oversight. The work identifies common risks such as specification drift, overreliance on generated outputs, and platform dependency, and empirically characterizes—for the first time—a structural trade-off between process depth and cross-agent portability, offering reproducible tools and a research agenda for future evaluation.

AI software developmentdevelopment frameworksLLM agents

Hot Scholars

EF

Eric Feron

Professor of Electrical Engineering
Control SystemsOperations ResearchComputer ScienceAerospace Engineering
ZX

Zhijie Xu

Pacific Northwest National Lab
Computational ModelingMultiscale ModelingNumerical Methods
ZL

Zherui Li

Beijing University of Posts and Telecommunications
Large Language ModelTrustworthy AILLM-based Agent
YW

Yucheng Wang

ETH Zürich
Multimodal LLMSpeech UnderstandingHuman-Computer Interaction