llm pipeline orchestration

Designs, implements, and evaluates orchestrated pipelines and workflows that coordinate large language models and LangChain components (chains, agents, retrievers, memory, connectors), external tools and data sources, and inference/deployment infrastructure. Builds LangChain-based integrations, applications, and frameworks and analyzes pipeline behavior (latency, cost, correctness, reliability) and component interactions to ensure robust end-to-end LLM application orchestration.

llmpipelineorchestration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing LLM agent systems suffer from tight coupling between logical workflows and underlying programming languages or deployment environments, resulting in high development/deployment complexity and poor maintainability. This paper proposes a declarative domain-specific language (DSL) tailored for LLM agent workflows—marking the first effort to universally abstract and unify common patterns such as RAG, API orchestration, and filtering. The DSL fully decouples workflow specification from execution semantics, enabling cross-language (Java/Python/Go) and cross-environment (cloud-native/on-premises) deployment. The system integrates multi-backend adapters, a lightweight workflow engine, and an automated metrics collection framework, natively supporting multi-strategy A/B testing and performance benchmarking. Evaluated in PayPal’s e-commerce setting, it reduces development time by 60% and accelerates deployment threefold; complex workflows shrink from >500 to <50 lines of code, achieve orchestration latency under 100 ms, and allow safe, non-engineer configuration.

Enables non-engineers to modify agent behaviors with low orchestration overheadSeparates agent workflow specification from implementation across languages and environmentsTransforms agent development from programming to configuration using a unified DSL

This study addresses the lack of systematic understanding regarding the practical usage patterns, reliability mechanisms, and autonomy levels of large language model (LLM) agents in low-code/no-code platforms. Drawing on over 6,000 publicly available n8n workflows, the authors employ large-scale data mining, structured log analysis, and qualitative coding to empirically characterize how LLM agents are deployed in real-world automation scenarios—specifically examining task distribution, workflow structure, tool invocation, and degrees of autonomy. The findings reveal that while LLMs are commonly embedded within complex workflows featuring control logic and human review steps, such workflows generally lack structured fault tolerance, repair loops, and approval mechanisms. Based on these insights, the study articulates ten empirical observations and five design implications to inform the development of more reliable and governable low-code platforms.

Agentic WorkflowsHuman-AI CollaborationLarge Language Models

This work addresses the challenges faced by large language model (LLM) agents operating over flat tool registries—namely, combinatorial explosion in decision space, context saturation, and degraded routing accuracy. To overcome these limitations, the authors propose a skill-tree-based hierarchical architecture that separates routing logic at internal nodes from execution at leaf nodes. Inspired by pushdown automata, the framework incorporates a LIFO stack-frame memory model and a lazy capability discovery mechanism, enabling isolated execution paths and scalable context management. The approach supports manifest-driven single-step execution loops and formal state modeling, significantly improving routing accuracy while reducing memory footprint and prompt costs under conditions of tool proliferation, multi-step workflows, and prompt exposure. This design meets enterprise-grade requirements for isolation and scalability.

context window saturationdecision-space explosionLLM agents

This study addresses the challenges of COBOL system modernization—namely, the scarcity of domain experts, the scale of legacy codebases, and stringent correctness requirements—and investigates the efficacy of large language model (LLM)-based orchestration strategies in automated COBOL-to-Python translation. For the first time, orchestration strategy is isolated as the sole variable within a unified experimental framework, enabling a direct comparison between deterministic and LLM-driven approaches. The findings reveal that deterministic orchestration achieves functional correctness on par with LLM-based control while significantly enhancing robustness, reducing inter-run performance variability, and cutting token consumption by up to 3.5×. These results demonstrate that deterministic orchestration offers superior stability and cost efficiency without compromising translation accuracy.

COBOL-to-Python modernizationdeterministic orchestrationexecution control

A Roadmap for Tamed Interactions with Large Language Models

Oct 28, 2025
VS
Vincenzo Scotti
🏛️ Karlsruhe Institute of Technology (KIT) | Ruhr Institute for Software Technology (paluno) | University of Duisburg-Essen

Large language models (LLMs) suffer from hallucination, unreliability, and uncontrolled behavior, hindering their trustworthy deployment in safety-critical workflows; existing reliability-enhancement tools are fragmented and lack a systematic framework. This paper introduces LSL (LLM Scripting Language), a domain-specific scripting language that embeds formal specifications, verifiable constraints, and explainability mechanisms directly into the LLM interaction process—enabling structured output constraints, programmable behavioral control, and decoupled execution governance. LSL unifies domain-specific language (DSL) design, formal verification, and runtime checking, significantly improving output reliability, consistency, and traceability. Experiments demonstrate that LSL effectively mitigates hallucination across diverse tasks, supports safe and controllable LLM integration, and establishes a novel interaction paradigm for trustworthy AI systems.

Addressing LLM unreliability and hallucination issuesDeveloping DSL to control LLM outputs and interactionsIntegrating verification and validation for trustworthy LLM applications

Latest Papers

What's happening recently
View more

This study addresses the growing challenge of runtime failures in locally deployed and fine-tuned open-source large language models (LLMs), which increasingly stem not from algorithmic flaws but from systemic fragility in the deployment stack. Through a large-scale empirical analysis of 705 real-world failure reports from the DeepSeek, Llama, and Qwen ecosystems, we identify and formally characterize three recurring phenomena: diagnostic divergence, system homogeneity, and lifecycle upgrade issues. By integrating fault log analysis, root cause categorization, and cross-ecosystem comparison, we construct the first reliability analysis framework for open-source LLM deployment, establishing the deployment stack as the primary source of failures. We further release the first public dataset of such failures, offering actionable insights to enhance deployment reliability.

deployment failuresLLM reliabilityopen-source LLMs

This work addresses the inadequacy of existing large language model (LLM) lifecycle frameworks, which predominantly emphasize operational efficiency while lacking explicit support for security-critical activities—such as data provenance, component signing, and access control—and failing to align governance requirements with specific lifecycle phases. The paper proposes the first security-oriented LLM system lifecycle model, structured not by workflow but by security boundaries, organizing 32 phases into four layered pipelines: data, model, distribution, and application, while integrating LLMOps and governance pillars. It uniquely identifies 13 distinct security-critical phases and exposes a structural imbalance wherein regulatory evidence is concentrated at deployment despite pivotal decisions occurring during development. By mapping key standards—including NIST AI RMF, the EU AI Act, and ISO/IEC 42001—the study establishes a phase-to-governance correspondence mechanism, yielding a comprehensive, lifecycle-spanning security analysis framework that offers structured guidance for compliance and secure design.

governance frameworklarge language modelsLLM systems

This work addresses the limitations of existing agent orchestration frameworks, which rely on external schedulers and incur substantial context overhead, require state-of-the-art large language models, and risk exposing proprietary workflows. To overcome these issues, the authors propose compiling multi-node agent workflows—comprising up to 55 nodes—directly into the weights of a small fine-tuned language model, thereby creating what they term “underground agents.” This approach provides the first systematic demonstration that complex workflows can be effectively internalized within model parameters. By integrating structured workflow representations, task-specific knowledge injection, and decision-hub modeling, the method achieves performance comparable to leading models on tasks such as travel booking, Zoom customer support, and insurance claims processing, while reducing inference costs by two orders of magnitude and substantially diminishing reliance on conventional orchestration frameworks.

Agent OrchestrationAgentic WorkflowsFine-tuned Models

This work addresses the inefficiency and unreliability of directly deploying large language models (LLMs) or their distilled variants for enterprise tasks, which are typically deterministic, structured, and heavily reliant on domain-specific knowledge under strict constraints of cost, latency, and reliability. To overcome these limitations, the authors propose a modular AI architecture that confines LLMs to structured information extraction while offloading knowledge storage and reasoning to dedicated knowledge bases and symbolic systems. This design circumvents the bottlenecks of monolithic models in terms of interpretability, reliability, and maintainability. Theoretical analysis and system implementation demonstrate that the proposed architecture offers a more efficient, transparent, and sustainable alternative to end-to-end LLM approaches, providing a scalable AI solution tailored for enterprise applications.

deterministic workflowsenterprise tasksknowledge dependency

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

Dec 18, 2025
HL
Hao Liang
🏛️ Peking University | Institute for Advanced Algorithms Research | OriginHub Technology | OpenDataLab | Shanghai Artificial Intelligence Laboratory | LLaMA-Factory Team

The era of large language models (LLMs) faces critical challenges including insufficient high-quality data supply, fragmented data preparation pipelines, poor reproducibility, and lack of model-in-the-loop support. Method: We propose the first LLM-driven, unified data preparation framework for data-centric AI, featuring system-level abstractions and PyTorch-style APIs for modular design. We introduce DataFlow-Agent—the first agent that synthesizes executable data pipelines end-to-end from natural language specifications—and integrate LLM-powered operator synthesis, iterative validation, 200+ reusable operators, and six domain-agnostic pipeline templates. Results: Experiments on Text-to-SQL, code generation, and mathematical reasoning show our synthesized data significantly outperforms human-annotated and domain-specific synthetic data. Remarkably, just 10K samples surpass the performance of models trained on the million-scale Infinity-Instruct dataset, empirically validating the decisive impact of data quality on model performance.

Addresses scalable, reliable data preparation for LLMsAutomates pipeline creation from natural language specificationsReplaces ad-hoc scripts with modular, reusable data transformations

Hot Scholars

LH

Lewei He

South China Normal University
3D PrintingDeep Learning
YT

Yuan Tian

Associate Professor, School of Computing, Queen's University, Canada
Data MiningSoftware EngineeringLLM for SEMachine Learning
WS

Wei Song

Nanjing University of Science and Technology
software engineeringsoftware analysisservice compositionprocess mining
YD

Ying Ding

Bill & Lewis Suit Professor, School of Information, Dell Med, University of Texas at Austin
AI in HealthKnowledge GraphScience of Science
HS

Houbing Song

IEEE Fellow, Co-EiC of TII, University of Maryland, Baltimore County
Neuro-symbolic AIAnomaly DetectionArtificial Intelligence of ThingsAutonomous and CPS