Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead of full-budget counterfactual evaluation in generative image communication by proposing the ACV-Gate framework. This framework integrates terminal value learning with adaptive candidate evaluation to establish a controllable computation allocation mechanism. Specifically, it employs an ensemble-aware student network, terminal advantage and regret training, and local minimum description length (MDL) with cost thresholding to selectively execute approximate evaluations or exact computations, thereby optimizing token selection and reconstruction quality. Experimental results on CIFAR-10 demonstrate that the proposed method improves PSNR by 0.636 dB while reducing the number of evaluations to 27.6% of those required by expert mode, significantly enhancing communication performance under low-bitrate conditions.
📝 Abstract
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated receiver-side reconstruction, resulting in substantial encoder-side computation. To address this problem, we propose ACV-Gate, an adaptive candidate evaluation framework that learns to approximate full-budget counterfactual evaluation and selectively assigns exact evaluations to the most informative candidates. Specifically, a set-aware student is trained using terminal advantages and regrets to predict candidate rankings directly, while a selective refinement mechanism evaluates only a bounded candidate set containing both Local-MDL and direct actions; cost-based thresholds further enable explicit control of the average evaluation workload. Experiments on CIFAR-10 show that ACV-Gate consistently improves reconstruction quality while substantially reducing candidate evaluations; at 0.20 bpp, the primary adaptive configuration improves PSNR over LocalMDL by 0.636 dB with only 2.13 candidate evaluations per image, corresponding to 27.60% of the calls required by the Exact-Full expert. Matched-candidate comparisons, synchronized GPU measurements, and evaluations on STL-10 and 384 *384 scale transfer further demonstrate consistent quality computation trade-offs, with particularly pronounced gains at low bit rates. These results show that combining terminal-value learning with selective candidate evaluation provides an effective and controllable mechanism for allocating encoder computation in packet-constrained generative image communication.
Problem

Research questions and friction points this paper is trying to address.

generative image communication
visual token selection
counterfactual reasoning
computational overhead
packet budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Image Communication
Counterfactual Reasoning
Adaptive Candidate Evaluation
Visual Token Selection
Computational Amortization
🔎 Similar Papers
Q
Qinglei Qi
School of Artificial Intelligence, Nanyang Normal University, Nanyang, China
Z
Zhihe Liang
School of Artificial Intelligence, Nanyang Normal University, Nanyang, China
F
Fengzhan Jing
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
S
Shenao Zhu
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
L
Lei Zhang
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
C
Chenyang Zhang
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
S
Shuqing He
School of Information Science and Engineering, Linyi University, Linyi 276000, China
J
Jia Guo
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China