Institution profile

Xiaohongshu

Industry researchasia · cn
Official website
Research library34linked papers
Opportunities0open roles
Selected work

Representative Papers

AdaptEvo: Adaptive Agent Learning with Evolving Supervision

Oct 08, 2026

This study addresses the challenges of imperfect supervision signals and dynamically evolving evaluation criteria in rule-driven decision-making by proposing a framework that integrates confidence-adaptive policy optimization with evolutionary decision knowledge. Methodologically, we design the CA-GRPO algorithm to dynamically balance outcome and process rewards, alongside an evolutionary module that distills reusable knowledge from failure cases. Experiments conducted on the Qwen3.6 backbone using a multimodal content moderation dataset demonstrate that the proposed framework significantly outperforms existing baselines on industrial-scale data. Furthermore, it exhibits exceptional robustness under shifting rule conditions, effectively overcoming the limitations inherent in fixed reward-mixing approaches.

0 citationsRead paper

Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval

Oct 08, 2026

This study addresses the limitation of fixed query representations in general-purpose multimodal retrieval, which hinders the exploitation of retrieved results to clarify information needs. To this end, we propose the Alternating Retrieval and Refinement (ARR) model, which introduces an autoregressive feedback mechanism that dynamically optimizes information need representations by alternating between retrieval and query embedding updates. During training, ARR combines stepwise contrastive supervised fine-tuning with reinforcement learning to precisely select informative feedback items, while employing a query-side adapter to enhance generalization. Experiments demonstrate that ARR significantly outperforms existing baselines on both in-domain and zero-shot benchmarks, validating that iterative feedback yields dual benefits for initial retrieval accuracy and downstream reasoning performance.

0 citationsRead paper

GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

Oct 06, 2026

This study addresses the issue wherein traditional distillation misidentifies critical steps due to trajectory drift in LLM agents operating under sparse rewards. To this end, we propose a graph-structure-enhanced online distillation method that constructs dependency graphs from environment states and integrates random walk scoring with divergence signals to precisely identify high-value steps. This work represents the first introduction of structural graph modeling into agent online distillation for causal credit assignment. Experimental results demonstrate that the proposed approach outperforms the strongest baseline by 5.8 percentage points on benchmarks such as ALFWorld, validating both the causal efficacy of structure-based credit assignment and its cross-domain transferability.

0 citationsRead paper

RubricArmor: Adversarial Evolution Improves LLM-Based Rubric Generation

Oct 04, 2026

This study addresses the vulnerability of large language model (LLM)-generated rubrics to reward hacking, which enables low-quality responses to receive undeservedly high scores. To mitigate this issue, we propose RubricArmor, an adversarial framework that pioneers the integration of an adversarial evolution mechanism during the rubric generation phase. By combining reinforcement learning with adversarial prompt engineering, the framework employs an iterative attack-simulation and repair-evolution algorithm to proactively identify and patch potential vulnerabilities, thereby achieving active defense at the generation stage rather than merely optimizing rubric granularity. Experimental results demonstrate that our approach significantly outperforms existing baselines and effectively enhances downstream rubric-based RLHF alignment performance.

0 citationsRead paper

Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation

Oct 01, 2026

This study addresses the high computational cost of teacher supervision in online knowledge distillation by proposing an adaptive supervision selection mechanism that leverages the student model's own successful trajectories as a reference. Specifically, this method precisely allocates teacher resources by filtering out failed samples, while integrating hidden state trajectory discrepancy analysis with a budget-aware token selection strategy to achieve efficient supervision. Experimental results demonstrate that the proposed approach maintains comparable inference performance while requiring only 3.46% to 5.02% of the teacher inputs, thereby substantially reducing the computational overhead associated with online distillation.

0 citationsRead paper
Recent publications

Latest Papers

AdaptEvo: Adaptive Agent Learning with Evolving Supervision

Oct 08, 2026

This study addresses the challenges of imperfect supervision signals and dynamically evolving evaluation criteria in rule-driven decision-making by proposing a framework that integrates confidence-adaptive policy optimization with evolutionary decision knowledge. Methodologically, we design the CA-GRPO algorithm to dynamically balance outcome and process rewards, alongside an evolutionary module that distills reusable knowledge from failure cases. Experiments conducted on the Qwen3.6 backbone using a multimodal content moderation dataset demonstrate that the proposed framework significantly outperforms existing baselines on industrial-scale data. Furthermore, it exhibits exceptional robustness under shifting rule conditions, effectively overcoming the limitations inherent in fixed reward-mixing approaches.

0 citationsRead paper

Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval

Oct 08, 2026

This study addresses the limitation of fixed query representations in general-purpose multimodal retrieval, which hinders the exploitation of retrieved results to clarify information needs. To this end, we propose the Alternating Retrieval and Refinement (ARR) model, which introduces an autoregressive feedback mechanism that dynamically optimizes information need representations by alternating between retrieval and query embedding updates. During training, ARR combines stepwise contrastive supervised fine-tuning with reinforcement learning to precisely select informative feedback items, while employing a query-side adapter to enhance generalization. Experiments demonstrate that ARR significantly outperforms existing baselines on both in-domain and zero-shot benchmarks, validating that iterative feedback yields dual benefits for initial retrieval accuracy and downstream reasoning performance.

0 citationsRead paper

GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

Oct 06, 2026

This study addresses the issue wherein traditional distillation misidentifies critical steps due to trajectory drift in LLM agents operating under sparse rewards. To this end, we propose a graph-structure-enhanced online distillation method that constructs dependency graphs from environment states and integrates random walk scoring with divergence signals to precisely identify high-value steps. This work represents the first introduction of structural graph modeling into agent online distillation for causal credit assignment. Experimental results demonstrate that the proposed approach outperforms the strongest baseline by 5.8 percentage points on benchmarks such as ALFWorld, validating both the causal efficacy of structure-based credit assignment and its cross-domain transferability.

0 citationsRead paper

RubricArmor: Adversarial Evolution Improves LLM-Based Rubric Generation

Oct 04, 2026

This study addresses the vulnerability of large language model (LLM)-generated rubrics to reward hacking, which enables low-quality responses to receive undeservedly high scores. To mitigate this issue, we propose RubricArmor, an adversarial framework that pioneers the integration of an adversarial evolution mechanism during the rubric generation phase. By combining reinforcement learning with adversarial prompt engineering, the framework employs an iterative attack-simulation and repair-evolution algorithm to proactively identify and patch potential vulnerabilities, thereby achieving active defense at the generation stage rather than merely optimizing rubric granularity. Experimental results demonstrate that our approach significantly outperforms existing baselines and effectively enhances downstream rubric-based RLHF alignment performance.

0 citationsRead paper

Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation

Oct 01, 2026

This study addresses the high computational cost of teacher supervision in online knowledge distillation by proposing an adaptive supervision selection mechanism that leverages the student model's own successful trajectories as a reference. Specifically, this method precisely allocates teacher resources by filtering out failed samples, while integrating hidden state trajectory discrepancy analysis with a budget-aware token selection strategy to achieve efficient supervision. Experimental results demonstrate that the proposed approach maintains comparable inference performance while requiring only 3.46% to 5.02% of the teacher inputs, thereby substantially reducing the computational overhead associated with online distillation.

0 citationsRead paper