Computer code validation via mixture model estimation
本文通过贝叶斯混合模型方法解决计算机代码验证问题,比较纯代码模型与偏差修正模型,并使用Metropolis-within-Gibbs算法进行推断。
本文通过贝叶斯混合模型方法解决计算机代码验证问题,比较纯代码模型与偏差修正模型,并使用Metropolis-within-Gibbs算法进行推断。
研究通过基于规则的逆生物合成方法,使用Qwen2.5-7B策略选择扩展分子,提高了在给定扩展次数下的解题率。
研究在非平稳环境下时间序列的早期分类问题,通过强化学习提出DQeND模型联合学习表示、分类和触发决策,优于传统分离设计方法。
This paper addresses core challenges in automatic text summarization evaluation—poor metric reproducibility, low correlation with human judgments, and the trade-off between computational cost and result stability. To this end, we introduce the first open-source, unified evaluation framework supporting standardized, reproducible comparison of diverse metrics, including ROUGE, G-Eval, SEval-Ex, and multiple LLM-based evaluators. Systematic experiments on SummEval reveal that LLM-based metrics exhibit substantial stochasticity and irreproducibility; quantitative analysis further confirms that high human alignment often comes at the cost of elevated computational overhead and reduced output stability. Our key contributions are: (1) the first unified framework compatible with heterogeneous evaluation paradigms; (2) the first empirical demonstration of reproducibility deficits in LLM-based evaluation; and (3) advancement of transparent, standardized evaluation protocols, establishing a reproducible benchmark for future summarization evaluation research.
Addressing the trade-off between performance and interpretability in summary evaluation, this paper proposes a fine-grained assessment framework based on atomic sentence decomposition: leveraging large language models (LLMs) to perform sentence-level extraction and semantic alignment between source documents and summaries, thereby constructing traceable and verifiable decision evidence chains. The method introduces the first sentence-level alignment evaluation paradigm, overcoming the limitations of conventional summary-level scoring. It is the first approach to achieve state-of-the-art (SOTA) performance while providing transparent, step-by-step reasoning paths and demonstrating robustness against hallucination. On the SummEval benchmark, it attains a consistency score of 0.580—significantly outperforming the GPT-4 evaluator (0.521)—thereby advancing both evaluation accuracy and interpretability.
本文通过贝叶斯混合模型方法解决计算机代码验证问题,比较纯代码模型与偏差修正模型,并使用Metropolis-within-Gibbs算法进行推断。
研究通过基于规则的逆生物合成方法,使用Qwen2.5-7B策略选择扩展分子,提高了在给定扩展次数下的解题率。
研究在非平稳环境下时间序列的早期分类问题,通过强化学习提出DQeND模型联合学习表示、分类和触发决策,优于传统分离设计方法。
This paper addresses core challenges in automatic text summarization evaluation—poor metric reproducibility, low correlation with human judgments, and the trade-off between computational cost and result stability. To this end, we introduce the first open-source, unified evaluation framework supporting standardized, reproducible comparison of diverse metrics, including ROUGE, G-Eval, SEval-Ex, and multiple LLM-based evaluators. Systematic experiments on SummEval reveal that LLM-based metrics exhibit substantial stochasticity and irreproducibility; quantitative analysis further confirms that high human alignment often comes at the cost of elevated computational overhead and reduced output stability. Our key contributions are: (1) the first unified framework compatible with heterogeneous evaluation paradigms; (2) the first empirical demonstration of reproducibility deficits in LLM-based evaluation; and (3) advancement of transparent, standardized evaluation protocols, establishing a reproducible benchmark for future summarization evaluation research.
Addressing the trade-off between performance and interpretability in summary evaluation, this paper proposes a fine-grained assessment framework based on atomic sentence decomposition: leveraging large language models (LLMs) to perform sentence-level extraction and semantic alignment between source documents and summaries, thereby constructing traceable and verifiable decision evidence chains. The method introduces the first sentence-level alignment evaluation paradigm, overcoming the limitations of conventional summary-level scoring. It is the first approach to achieve state-of-the-art (SOTA) performance while providing transparent, step-by-step reasoning paths and demonstrating robustness against hallucination. On the SummEval benchmark, it attains a consistency score of 0.580—significantly outperforming the GPT-4 evaluator (0.521)—thereby advancing both evaluation accuracy and interpretability.