Orbital Detection: On Maximum-Entropy Priors
该论文提出使用最大熵先验方法降低软输入检测的计算成本,通过引入轨道先验分布,将每符号计算复杂度从O(M)降至O(L),同时保持了良好的误码率性能。
该论文提出使用最大熵先验方法降低软输入检测的计算成本,通过引入轨道先验分布,将每符号计算复杂度从O(M)降至O(L),同时保持了良好的误码率性能。
研究通过优化SKILL文件来提高编码代理性能,使用合并的拉取请求作为更难任务,并基于代理表现评分,发现GEPA方法有效提升了文档质量。
该文介绍了一种使用大型语言模型辅助开发静态验证软件的工具Eiffel-tools,通过与语言服务器协议结合,并利用形式化验证器提高代码生成和修正的准确性。
本文针对异构敏捷地球观测卫星调度问题,提出了一种结合强化学习引导的进化策略优化框架,通过解码器和基于群体的搜索方法实现高效优化。
This study addresses the limitations of holistic scoring in metaphor explanation evaluation, which often overlooks quality structure and human disagreement. We propose a cognition-driven, six-dimensional assessment framework to capture these nuances. Through large-scale annotation and clustering analysis, we reveal the multidimensionality of explanation quality and systematic patterns of disagreement, validating that an automated evaluation pipeline can effectively recover this structure. Our results demonstrate that automatic models can predict key dimensions, with prediction errors significantly correlating with human disagreement. By overcoming the constraints of single-score metrics, this work establishes a fine-grained, diagnostically valuable evaluation paradigm for open-ended generation tasks, offering deeper insights into model performance and human alignment.
该论文提出使用最大熵先验方法降低软输入检测的计算成本,通过引入轨道先验分布,将每符号计算复杂度从O(M)降至O(L),同时保持了良好的误码率性能。
研究通过优化SKILL文件来提高编码代理性能,使用合并的拉取请求作为更难任务,并基于代理表现评分,发现GEPA方法有效提升了文档质量。
该文介绍了一种使用大型语言模型辅助开发静态验证软件的工具Eiffel-tools,通过与语言服务器协议结合,并利用形式化验证器提高代码生成和修正的准确性。
本文针对异构敏捷地球观测卫星调度问题,提出了一种结合强化学习引导的进化策略优化框架,通过解码器和基于群体的搜索方法实现高效优化。
This study addresses the limitations of holistic scoring in metaphor explanation evaluation, which often overlooks quality structure and human disagreement. We propose a cognition-driven, six-dimensional assessment framework to capture these nuances. Through large-scale annotation and clustering analysis, we reveal the multidimensionality of explanation quality and systematic patterns of disagreement, validating that an automated evaluation pipeline can effectively recover this structure. Our results demonstrate that automatic models can predict key dimensions, with prediction errors significantly correlating with human disagreement. By overcoming the constraints of single-score metrics, this work establishes a fine-grained, diagnostically valuable evaluation paradigm for open-ended generation tasks, offering deeper insights into model performance and human alignment.