integrate decoding pipelines

Designs, builds, and analyzes end-to-end decoding systems that integrate multiple decoding components and stages into deployable pipelines, including orchestration, batching, and performance-aware integration. Implements and integrates mechanisms for controlled decoding (constraints, prefixes, steering, or conditional generation) and the modules that convert intermediate model outputs into final signals, for example vocoder design for waveform synthesis.

integratedecodingpipelines

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the frequent failures of large language models in multi-step tool orchestration tasks, often caused by incorrect API invocation sequences or improper parameter passing. To tackle this, the authors construct a reinforcement learning environment grounded in cached real-world API responses and propose a constrained data synthesis approach to generate multi-step execution trajectories with controllable complexity. They further design a hierarchical reward mechanism that separately evaluates the atomic validity of individual API calls and the logical correctness of the overall orchestration. This framework overcomes the limitations of traditional binary rewards and purely simulated data, achieving significant improvements in episode-level accuracy on the ComplexFuncBench benchmark. Ablation studies confirm the necessity of both reward signals for effective learning.

API dependencyLLM trainingmulti-step tool orchestration

Existing runtime harnesses for programming agents suffer from either oversimplification or excessive complexity, lacking a clear and concise architectural paradigm. This work proposes a harness design centered on the request lifecycle, explicitly delineating three core boundaries: model, execution, and state. By orchestrating a structured sequence—comprising context construction, model decision-making, environmental action, observation feedback, and state continuation—the design enables cross-request state persistence and continual self-improvement through bootstrapping. We implement this paradigm in Coderlet, an open-source prototype system, demonstrating its efficacy in coordinating code generation, environment interaction, and state management. The resulting framework provides a scalable foundation for building high-performance programming agents.

harnessmodel-environment interactionprogramming agent

Multi-objective controllable generation faces challenges including diverse user requirements, inefficient parameter-level control, high decoding-guidance overhead, and overreliance on single-model capabilities. Method: This paper proposes MAGE, a two-stage framework that (i) identifies and resolves compatibility mismatches between guidance and base models; (ii) dynamically constructs multi-objective fused base models via model merging and unifies explicit/implicit value models as collaborative guidance agents; and (iii) integrates linear mode connectivity analysis, predictive ensembling, and two-stage guided decoding to jointly optimize parameter- and decoding-level control. Contribution/Results: Experiments demonstrate that MAGE significantly outperforms state-of-the-art methods in controllability, Pareto optimality, and cross-task adaptability, while reducing memory overhead and enhancing multi-objective coordination.

Addressing insufficient control in multi-objective text generationReducing space overhead from multiple expert model aggregationResolving compatibility between guidance and base models

This work proposes Code-Flow, a multi-stage training paradigm designed to model the dynamic evolution of code in software development and enhance large language models’ capabilities in intelligent programming, agent-based software engineering, and complex tool invocation. The approach integrates pretraining, an intermediate training phase grounded in agent execution trajectories, and a bifurcated post-training strategy comprising a reasoning-driven reinforcement learning path (Thinking) and a general-purpose instruction-tuning path (Instruct). A Loop architecture is introduced to balance performance gains with deployment overhead. Trained with extended context windows of 32k and 128k tokens, the resulting IQuest-Coder-V1 series achieves state-of-the-art performance on critical benchmarks spanning agent-driven software engineering, competitive programming, and sophisticated tool usage.

agentic software engineeringcode intelligencedynamic code evolution

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic understanding in co-designing the prefill and decode phases for heterogeneous large language model inference, which hinders deployment efficiency. Focusing on four key design axes—accelerator architecture, precision, interconnect, and KV residency—the study reveals that only a subset of these factors forms strongly coupled constraints and identifies three pivotal boundary decisions: compute placement, KV representation, and KV ownership. Integrating empirical insights from industrial deployments with runtime analysis, the authors propose a runtime role–driven precision strategy, a byte-level KV transfer mechanism, and explicit ownership management to construct a multi-axis design space model. The resulting framework yields actionable design guidelines that substantially improve heterogeneous inference efficiency.

design spaceheterogeneous LLM inferenceKV state management

This work addresses the challenge of harmonizing explicit user commands with implicit interaction context in expressive speech synthesis for voice assistants. We propose Harness TTS, which formulates style control as a closed-set prompt-tool routing task grounded in large language models (LLMs). By introducing a structured prompt-tool registry, a priority-aware observation mechanism, and a lightweight control layer, our approach enables externalized, auditable, and low-latency context-aware modulation of TTS expressive behavior. Integrated with a Qwen3-4B planner and CosyVoice3/VoxCPM2 synthesizers, the system achieves Top-1 routing accuracies of 74.3%, 43.0%, and 64.6% in explicit, implicit, and conflicting scenarios, respectively. Synthesized speech significantly outperforms instruction-only control, yielding UTMOSv2 gains of 0.11–0.38 while maintaining first-token latency under 50 ms.

context-awareexpressive speech synthesisstyle control

This study addresses the challenge of designing effective inference-time harnesses to improve the long-term execution success of large language model agents in complex tasks. Recognizing that excessive decomposition or guidance can degrade performance, the work formalizes the harness design as a trajectory alignment problem and decouples it into two mechanisms: task decomposition and guided execution. It systematically investigates the impact of workflow granularity, retry budgets, and action reweighting. The authors propose a “partial harness” strategy—specifying only the initial steps—which effectively steers execution while avoiding failure modes such as over-decomposition, over-pruning, and hallucinated actions. Empirical validation in both synthetic environments and real-world terminal agent tasks demonstrates that this approach significantly enhances task completion rates, outperforming fully structured workflows.

execution trajectoriesguided executionharness design

Traditional AI systems rely on fixed monolithic models, which struggle to dynamically allocate resources, decompose tasks, or update knowledge in response to varying inputs, leading to degraded performance and increased costs. This work proposes the first system-level design methodology for distributed composite AI systems, formulating a design space through workflow topologies and configuration choices and identifying eight core design patterns. The framework jointly optimizes model selection and runtime parameters, enabling task decomposition, multi-model orchestration, and explicit control logic, thereby facilitating a shift from static monolithic architectures toward dynamic, composable, and adaptive ones. Evaluated across three case studies, the approach reduces latency by up to 60% and cost by up to 71%, with only a 2.5–4 percentage point drop in accuracy.

Compound AI SystemsDistributed AIModel-Centric Design

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
JL

Jian Luan

Toshiba, Microsoft, Xiaomi
LLMVLMTTSSinging Synthesis
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
SC

Souradip Chakraborty

University of Maryland, College Park | Past : ML Research@Walmart Labs
Reinforcement LearningDeep LearningRobustnessUncertainty