Institution profile

Tianjin Normal University

Academic institutionasia · cn
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training

Sep 27, 2026

This study addresses the memory waste and recomputation overhead caused by fixed checkpointing strategies in GRPO training by proposing HiLoRe, an adaptive memory management method. For the first time, HiLoRe integrates the analytical update structure of GRPO into state fidelity allocation. By modeling policy update exposure and predicting approximation risk via loss coefficients, it dynamically allocates resources across high-precision storage, low-precision compression, and deterministic recomputation, thereby achieving memory optimization under risk-budget calibration. Experiments demonstrate that under memory-constrained conditions, HiLoRe improves Actor update throughput by 13.5% compared to gradient checkpointing (GC), while maintaining downstream task performance within a 0.6 percentage point margin. These results confirm that the proposed approach effectively balances training efficiency with model quality.

0 citationsRead paper

Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication

Sep 24, 2026

This study addresses the high computational overhead of full-budget counterfactual evaluation in generative image communication by proposing the ACV-Gate framework. This framework integrates terminal value learning with adaptive candidate evaluation to establish a controllable computation allocation mechanism. Specifically, it employs an ensemble-aware student network, terminal advantage and regret training, and local minimum description length (MDL) with cost thresholding to selectively execute approximate evaluations or exact computations, thereby optimizing token selection and reconstruction quality. Experimental results on CIFAR-10 demonstrate that the proposed method improves PSNR by 0.636 dB while reducing the number of evaluations to 27.6% of those required by expert mode, significantly enhancing communication performance under low-bitrate conditions.

0 citationsRead paper

Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

Aug 17, 2026

This study addresses the misalignment between existing visual token selection criteria and reconstruction quality under fixed bandwidth constraints. We propose Gated Counterfactual Rectification (GCR-C), a method that constructs candidate sets and performs full-budget counterfactual evaluations to dynamically replace baseline actions only when positive gains are confirmed. This approach effectively bridges the gap between selection strategies and final reconstruction outcomes. Experiments demonstrate that GCR-C significantly improves reconstruction quality at low-to-medium bitrates across diverse datasets and channel conditions without increasing actual bitrate consumption. Furthermore, the method exhibits robust generalization capabilities, establishing a novel paradigm for communication-aware reconstruction tasks.

0 citationsRead paper
Recent publications

Latest Papers

HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training

Sep 27, 2026

This study addresses the memory waste and recomputation overhead caused by fixed checkpointing strategies in GRPO training by proposing HiLoRe, an adaptive memory management method. For the first time, HiLoRe integrates the analytical update structure of GRPO into state fidelity allocation. By modeling policy update exposure and predicting approximation risk via loss coefficients, it dynamically allocates resources across high-precision storage, low-precision compression, and deterministic recomputation, thereby achieving memory optimization under risk-budget calibration. Experiments demonstrate that under memory-constrained conditions, HiLoRe improves Actor update throughput by 13.5% compared to gradient checkpointing (GC), while maintaining downstream task performance within a 0.6 percentage point margin. These results confirm that the proposed approach effectively balances training efficiency with model quality.

0 citationsRead paper

Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication

Sep 24, 2026

This study addresses the high computational overhead of full-budget counterfactual evaluation in generative image communication by proposing the ACV-Gate framework. This framework integrates terminal value learning with adaptive candidate evaluation to establish a controllable computation allocation mechanism. Specifically, it employs an ensemble-aware student network, terminal advantage and regret training, and local minimum description length (MDL) with cost thresholding to selectively execute approximate evaluations or exact computations, thereby optimizing token selection and reconstruction quality. Experimental results on CIFAR-10 demonstrate that the proposed method improves PSNR by 0.636 dB while reducing the number of evaluations to 27.6% of those required by expert mode, significantly enhancing communication performance under low-bitrate conditions.

0 citationsRead paper

Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

Aug 17, 2026

This study addresses the misalignment between existing visual token selection criteria and reconstruction quality under fixed bandwidth constraints. We propose Gated Counterfactual Rectification (GCR-C), a method that constructs candidate sets and performs full-budget counterfactual evaluations to dynamically replace baseline actions only when positive gains are confirmed. This approach effectively bridges the gap between selection strategies and final reconstruction outcomes. Experiments demonstrate that GCR-C significantly improves reconstruction quality at low-to-medium bitrates across diverse datasets and channel conditions without increasing actual bitrate consumption. Furthermore, the method exhibits robust generalization capabilities, establishing a novel paradigm for communication-aware reconstruction tasks.

0 citationsRead paper