VERA: Verifiable Feasibility Representations with Counterfactual Credit for Constrained Multi-Agent Control

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the execution failures in constrained multi-agent control caused by the decoupling of action feasibility prediction and credit assignment. To overcome this, we propose a centralized training with decentralized execution (CTDE) framework. Methodologically, we pioneer the use of action-conditioned feasibility as an auditable training interface, constructing a five-dimensional verifiable feasibility representation. Furthermore, we introduce a counterfactual group relative advantage algorithm to achieve geometry-dependent, precise credit assignment. Experimental results demonstrate that the proposed method improves the success rate by 8.74% and reduces the violation rate by 52.19% compared to baselines in Space-Air-Ground Integrated Network (SAGIN) environments, while exhibiting superior generalization capability and low computational overhead.
📝 Abstract
Constrained multi-agent control requires more than predicting rewarding actions: an action can cease to be executable as contact windows, shared capacity, and deadlines change. We introduce VERA, a centralized-training, decentralized-execution framework that separates feasibility estimation from credit assignment. Each actor predicts a five-dimensional verifiable feasibility representation (VFR). After an action is proposed, exact action-conditioned margins available only during training supervise that representation, while a counterfactual group-relative advantage (CGRA) ranks candidate representation-action pairs. Execution uses one actor pass and no privileged state. In a dynamic space-air-ground integrated network (SAGIN), VERA obtains 55.33% +/- 3.60% success with 0.45% +/- 0.81% coverage violation, within 1.33 percentage points of a privileged-mask reference. With rewards matched over ten paired seeds, VERA improves success over the strongest baseline by 8.74 percentage points (p=0.023) and reduces violation by 52.19 percentage points (p=5.7e-8). A ten-seed 4-by-2 factorial attributes a 14.16-16.48 percentage-point gain to CGRA across handcrafted, learned, random, and latent representations; evaluation on seven unseen topologies preserves a 24.33-30.02 percentage-point advantage over multi-agent proximal policy optimization. From 10 to 40 users, success remains 50.1-53.8%, and VFR adds only 0.026 ms to a central processing unit (CPU) actor step. Cross-domain tests further identify the governing condition: counterfactual credit succeeds when candidate scores respect shared constraints and fails under incompatible reward geometries. These results establish action-conditioned feasibility as an auditable training interface and counterfactual credit as a geometry-dependent optimization mechanism.
Problem

Research questions and friction points this paper is trying to address.

constrained multi-agent control
action feasibility
dynamic constraints
credit assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifiable Feasibility Representation
Counterfactual Credit Assignment
Constrained Multi-Agent Control
Centralized Training Decentralized Execution
Space-Air-Ground Integrated Network
🔎 Similar Papers
No similar papers found.