Institution profile

Beijing Foreign Studies University

Academic institutionasia · cn
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Sep 30, 2026

This study addresses the signal imbalance in multi-reward reinforcement learning caused by uneven reward activation densities, which constrains the effectiveness of multi-objective training for large language models. To this end, it proposes a density-aware reward aggregation mechanism that reveals the intrinsic relationship between advantage energy and activation density. Specifically, the method dynamically adjusts the weights of sparse rewards through inverse square root density correction. Furthermore, by integrating GDPO normalization, it introduces an adaptive weighting algorithm that requires no modifications to the underlying objectives. Experimental results demonstrate that this approach reduces the number of training steps by 26% on tool-calling tasks and significantly improves length compliance in mathematical reasoning, while maintaining competitive overall performance.

0 citationsRead paper

QSCP: Beyond Class-Name Prompts for Query-Guided Semantic Change Parsing

Sep 27, 2026

This study addresses the limitation of existing remote sensing change detection methods that rely solely on category prompts and cannot interpret natural language queries. To overcome this, we propose a query-guided semantic change parsing framework that integrates intent parsing, semantic slot filling, bidirectional visual evidence composition, and a query-conditioned decoder. Moving beyond traditional object-class localization paradigms, our approach enables explicit reasoning over source and target states, supports synonymous expressions and complex instructions, and outputs specific masks alongside bi-temporal semantic maps. Experimental results demonstrate that the proposed framework outperforms RCDNet on the SECOND dataset, significantly improving end-to-end semantic prediction accuracy. Furthermore, evaluations on the WHU-CDC dataset validate its robust cross-domain transferability.

0 citationsRead paper
Recent publications

Latest Papers

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Sep 30, 2026

This study addresses the signal imbalance in multi-reward reinforcement learning caused by uneven reward activation densities, which constrains the effectiveness of multi-objective training for large language models. To this end, it proposes a density-aware reward aggregation mechanism that reveals the intrinsic relationship between advantage energy and activation density. Specifically, the method dynamically adjusts the weights of sparse rewards through inverse square root density correction. Furthermore, by integrating GDPO normalization, it introduces an adaptive weighting algorithm that requires no modifications to the underlying objectives. Experimental results demonstrate that this approach reduces the number of training steps by 26% on tool-calling tasks and significantly improves length compliance in mathematical reasoning, while maintaining competitive overall performance.

0 citationsRead paper

QSCP: Beyond Class-Name Prompts for Query-Guided Semantic Change Parsing

Sep 27, 2026

This study addresses the limitation of existing remote sensing change detection methods that rely solely on category prompts and cannot interpret natural language queries. To overcome this, we propose a query-guided semantic change parsing framework that integrates intent parsing, semantic slot filling, bidirectional visual evidence composition, and a query-conditioned decoder. Moving beyond traditional object-class localization paradigms, our approach enables explicit reasoning over source and target states, supports synonymous expressions and complex instructions, and outputs specific masks alongside bi-temporal semantic maps. Experimental results demonstrate that the proposed framework outperforms RCDNet on the SECOND dataset, significantly improving end-to-end semantic prediction accuracy. Furthermore, evaluations on the WHU-CDC dataset validate its robust cross-domain transferability.

0 citationsRead paper