Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational inefficiency and risk of overriding correct intermediate results caused by blindly executing all components in multi-agent LLM workflows. To this end, we propose LW2S, a framework that introduces the first counterfactual credit assignment mechanism based on controlled skip interventions. By formulating component omission as a counterfactual problem, LW2S learns action-specific safety models to dynamically determine which steps can be safely skipped. Combined with holdout set calibration and domain-native guards, the framework enables adaptive component omission. Experimental results demonstrate that LW2S significantly reduces token overhead on tasks such as mathematical reasoning while maintaining or even improving overall accuracy, effectively validating the optimization value of conditioning on component-level utility.
📝 Abstract
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introduce Learning What to Skip (LW2S), which learns action-specific safety models from these interventions and combines held-out calibration with domain-native guards to select skips. When an early skip is rejected, the controller can continue execution and reconsider a later component. Across mathematical reasoning, multiple-choice QA, and code generation with two instruction-model families, LW2S reduces recorded token cost while matching or improving aggregate full-workflow accuracy in the evaluated settings. Scale-up and second-topology experiments further examine component redundancy, while shared-error cases reveal why agreement alone is insufficient for skip selection. These findings connect efficient workflow execution to learning the conditional utility of individual components.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent LLM Workflows
Counterfactual Credit Assignment
Component Omission
Computational Efficiency
Token Cost Reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Credit Assignment
Multi-Agent LLM Workflows
Learning What to Skip (LW2S)
Component Omission
Efficient Workflow Execution