Institution profile

SGIT AI

Research institution
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLMs with Evolution Strategies

Oct 04, 2026

This study addresses the limitations of Evolution Strategies (ES) in fine-tuning large language models, specifically their inability to dynamically perceive problem difficulty and the efficiency degradation caused by feedback compression. To overcome these challenges, this work proposes the DART-ES framework, which constructs a dynamic difficulty state by estimating local solvability without requiring auxiliary difficulty models or backpropagation. This mechanism jointly guides continuous difficulty reweighting and targeted replay of rare samples, thereby optimizing perturbation evaluation and data allocation. Evaluated on tasks such as mathematical reasoning, DART-ES significantly outperforms standard ES and achieves performance comparable to GRPO while substantially reducing runtime and memory overhead. These results establish DART-ES as a promising new paradigm for efficient full-parameter fine-tuning.

0 citationsRead paper
Recent publications

Latest Papers

DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLMs with Evolution Strategies

Oct 04, 2026

This study addresses the limitations of Evolution Strategies (ES) in fine-tuning large language models, specifically their inability to dynamically perceive problem difficulty and the efficiency degradation caused by feedback compression. To overcome these challenges, this work proposes the DART-ES framework, which constructs a dynamic difficulty state by estimating local solvability without requiring auxiliary difficulty models or backpropagation. This mechanism jointly guides continuous difficulty reweighting and targeted replay of rare samples, thereby optimizing perturbation evaluation and data allocation. Evaluated on tasks such as mathematical reasoning, DART-ES significantly outperforms standard ES and achieves performance comparable to GRPO while substantially reducing runtime and memory overhead. These results establish DART-ES as a promising new paradigm for efficient full-parameter fine-tuning.

0 citationsRead paper