Who Aggregates Information? Screening, Rent, and the Coexistence of CLOB and AMM Prediction Markets
研究通过引入尾部需求模型,分析了CLOB和LMSR在预测市场中共存的现象,解释了信息聚合机制及不同市场模式下的收益来源。
研究通过引入尾部需求模型,分析了CLOB和LMSR在预测市场中共存的现象,解释了信息聚合机制及不同市场模式下的收益来源。
为解决开源语言模型的可复现性问题,通过确定训练中的非确定性来源顺序,实现跨硬件的独立复现,引入完全可审计的透明度层级。
The design mechanism of feedback context in self-distillation remains unclear. This work proposes a step-aligned critique feedback mechanism that provides structured guidance precisely at points of reasoning errors while avoiding interference with correct behaviors, thereby enhancing learning efficiency. Within a self-distillation framework, the study systematically compares three feedback forms: binary rewards (GRPO), reference solutions, and step-aligned critiques, complemented by per-token advantage analysis to evaluate their effectiveness. Experimental results demonstrate that step-aligned critique improves performance by 16.11 points over GRPO and by 5.27 points over reference solutions on the Avg@12 metric, confirming the critical importance of aligning feedback with the structural trajectory of reasoning.
Existing large language model (LLM) routing approaches suffer from limitations in cost efficiency or training overhead. This work proposes a lightweight linear routing mechanism based on ridge regression, which, to the best of our knowledge, is the first to apply linear models to dynamic selection among multi-domain expert LLMs. By leveraging input features derived from causal language modeling and reasoning tasks, the method enables low-overhead, highly generalizable routing decisions and supports dynamic addition or removal of expert models without retraining the router. Experimental results demonstrate that the approach matches baseline performance across two task categories and significantly outperforms existing methods in reasoning tasks, achieving 98.4% normalized performance.
This work addresses the limitations of traditional homogeneous parallel search, which is constrained by the inductive bias of a single language model and struggles to generate behavioral novelty. The authors propose the DEI framework, which for the first time employs heterogeneous large language models as mutation operators within distributed evolutionary nodes. By leveraging non-blocking collective communication to share local optima, DEI establishes a cross-model adversarial-cooperative mechanism that enhances both diversity and robustness. Empirical results demonstrate that model heterogeneity—not merely parallel scale—is the key driver of improved performance in language model–based quality-diversity (LLM-QD) optimization. On the Core War benchmark, a four-node heterogeneous system achieves a 124% increase in QD-Score (45.90 vs. 20.46) and a 28% improvement in coverage (80.6% vs. 63.0%) over the single-node baseline, consistently outperforming homogeneous parallel alternatives.
研究通过引入尾部需求模型,分析了CLOB和LMSR在预测市场中共存的现象,解释了信息聚合机制及不同市场模式下的收益来源。
为解决开源语言模型的可复现性问题,通过确定训练中的非确定性来源顺序,实现跨硬件的独立复现,引入完全可审计的透明度层级。
The design mechanism of feedback context in self-distillation remains unclear. This work proposes a step-aligned critique feedback mechanism that provides structured guidance precisely at points of reasoning errors while avoiding interference with correct behaviors, thereby enhancing learning efficiency. Within a self-distillation framework, the study systematically compares three feedback forms: binary rewards (GRPO), reference solutions, and step-aligned critiques, complemented by per-token advantage analysis to evaluate their effectiveness. Experimental results demonstrate that step-aligned critique improves performance by 16.11 points over GRPO and by 5.27 points over reference solutions on the Avg@12 metric, confirming the critical importance of aligning feedback with the structural trajectory of reasoning.
Existing large language model (LLM) routing approaches suffer from limitations in cost efficiency or training overhead. This work proposes a lightweight linear routing mechanism based on ridge regression, which, to the best of our knowledge, is the first to apply linear models to dynamic selection among multi-domain expert LLMs. By leveraging input features derived from causal language modeling and reasoning tasks, the method enables low-overhead, highly generalizable routing decisions and supports dynamic addition or removal of expert models without retraining the router. Experimental results demonstrate that the approach matches baseline performance across two task categories and significantly outperforms existing methods in reasoning tasks, achieving 98.4% normalized performance.
This work addresses the limitations of traditional homogeneous parallel search, which is constrained by the inductive bias of a single language model and struggles to generate behavioral novelty. The authors propose the DEI framework, which for the first time employs heterogeneous large language models as mutation operators within distributed evolutionary nodes. By leveraging non-blocking collective communication to share local optima, DEI establishes a cross-model adversarial-cooperative mechanism that enhances both diversity and robustness. Empirical results demonstrate that model heterogeneity—not merely parallel scale—is the key driver of improved performance in language model–based quality-diversity (LLM-QD) optimization. On the Core War benchmark, a four-node heterogeneous system achieves a 124% increase in QD-Score (45.90 vs. 20.46) and a 28% improvement in coverage (80.6% vs. 63.0%) over the single-node baseline, consistently outperforming homogeneous parallel alternatives.