Exact Quantile Balancing and Load-Error Injection for Mixture-of-Experts
本文通过引入精确分位数平衡和负载误差注入方法,解决了Mixture-of-Experts训练中的全局与局部负载不平衡问题。
本文通过引入精确分位数平衡和负载误差注入方法,解决了Mixture-of-Experts训练中的全局与局部负载不平衡问题。
研究解决了高质量预训练检查点在后续训练中表现不佳的问题,通过分析30B混合专家模型训练流程中的检查点质量,发现高解决方案密度的检查点更优。
This work addresses the susceptibility of conditional log-probability–based scoring in multiple-choice evaluation to answer length bias, which can unfairly penalize or overfavor longer responses. To mitigate this issue, the paper proposes Bayesian Accuracy—a plug-and-play scoring method grounded in Bayesian posterior probabilities. By explicitly modeling the prior distribution over answer lengths, the approach effectively eliminates linear length bias without requiring additional forward passes. Evaluated across diverse benchmarks and few-shot settings, Bayesian Accuracy substantially reduces empirical length bias and enhances the fairness and reliability of model assessment compared to standard accuracy and length-normalization baselines.
Existing game-theoretic models in adversarial domains such as law often overlook language as a mechanism of persuasion, failing to capture the nuanced, discourse-driven nature of strategic interaction. This work proposes a Strategic Courtroom Framework that treats language as a first-class strategic action space. It introduces a heterogeneous multi-agent system grounded in nine interpretable personality traits and incorporates a reinforcement learning–driven dynamic trait orchestrator to generate adaptive persuasive strategies tailored to opponents and case specifics. Evaluated using DeepSeek-R1 and Gemini 2.5 Pro across 10 synthetic cases, 84 three-trait combinations, and over 7,000 simulated trials, the framework demonstrates that heterogeneous agent teams significantly outperform homogeneous ones, and dynamically composed traits surpass handcrafted strategies, with quantitative and charismatic traits contributing most prominently to persuasive efficacy.
This work addresses the lack of systematic investigation into format selection and performance trade-offs in existing low-bit quantization-aware training (QAT) methods, as well as their insufficient evaluation on generative tasks. To this end, we propose the first integration of k-means clustering into QAT for 1-bit weight quantization, optimizing generative performance under a fixed inference memory budget. Our approach transcends the limitations of conventional integer-based quantization schemes by leveraging learned cluster centroids to better preserve model fidelity at ultra-low bitwidths. Experimental results demonstrate that, under identical memory constraints, our method significantly outperforms state-of-the-art integer quantization approaches while maintaining compatibility with general-purpose hardware for efficient deployment.
本文通过引入精确分位数平衡和负载误差注入方法,解决了Mixture-of-Experts训练中的全局与局部负载不平衡问题。
研究解决了高质量预训练检查点在后续训练中表现不佳的问题,通过分析30B混合专家模型训练流程中的检查点质量,发现高解决方案密度的检查点更优。
This work addresses the susceptibility of conditional log-probability–based scoring in multiple-choice evaluation to answer length bias, which can unfairly penalize or overfavor longer responses. To mitigate this issue, the paper proposes Bayesian Accuracy—a plug-and-play scoring method grounded in Bayesian posterior probabilities. By explicitly modeling the prior distribution over answer lengths, the approach effectively eliminates linear length bias without requiring additional forward passes. Evaluated across diverse benchmarks and few-shot settings, Bayesian Accuracy substantially reduces empirical length bias and enhances the fairness and reliability of model assessment compared to standard accuracy and length-normalization baselines.
Existing game-theoretic models in adversarial domains such as law often overlook language as a mechanism of persuasion, failing to capture the nuanced, discourse-driven nature of strategic interaction. This work proposes a Strategic Courtroom Framework that treats language as a first-class strategic action space. It introduces a heterogeneous multi-agent system grounded in nine interpretable personality traits and incorporates a reinforcement learning–driven dynamic trait orchestrator to generate adaptive persuasive strategies tailored to opponents and case specifics. Evaluated using DeepSeek-R1 and Gemini 2.5 Pro across 10 synthetic cases, 84 three-trait combinations, and over 7,000 simulated trials, the framework demonstrates that heterogeneous agent teams significantly outperform homogeneous ones, and dynamically composed traits surpass handcrafted strategies, with quantitative and charismatic traits contributing most prominently to persuasive efficacy.
This work addresses the lack of systematic investigation into format selection and performance trade-offs in existing low-bit quantization-aware training (QAT) methods, as well as their insufficient evaluation on generative tasks. To this end, we propose the first integration of k-means clustering into QAT for 1-bit weight quantization, optimizing generative performance under a fixed inference memory budget. Our approach transcends the limitations of conventional integer-based quantization schemes by leveraging learned cluster centroids to better preserve model fidelity at ultra-low bitwidths. Experimental results demonstrate that, under identical memory constraints, our method significantly outperforms state-of-the-art integer quantization approaches while maintaining compatibility with general-purpose hardware for efficient deployment.