Institution profile

AgroParisTech

Academic institutioneurope · fr
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume

Aug 29, 2025

This paper addresses core challenges in automatic text summarization evaluation—poor metric reproducibility, low correlation with human judgments, and the trade-off between computational cost and result stability. To this end, we introduce the first open-source, unified evaluation framework supporting standardized, reproducible comparison of diverse metrics, including ROUGE, G-Eval, SEval-Ex, and multiple LLM-based evaluators. Systematic experiments on SummEval reveal that LLM-based metrics exhibit substantial stochasticity and irreproducibility; quantitative analysis further confirms that high human alignment often comes at the cost of elevated computational overhead and reduced output stability. Our key contributions are: (1) the first unified framework compatible with heterogeneous evaluation paradigms; (2) the first empirical demonstration of reproducibility deficits in LLM-based evaluation; and (3) advancement of transparent, standardized evaluation protocols, establishing a reproducible benchmark for future summarization evaluation research.

0 citationsRead paper

SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation

May 04, 2025

Addressing the trade-off between performance and interpretability in summary evaluation, this paper proposes a fine-grained assessment framework based on atomic sentence decomposition: leveraging large language models (LLMs) to perform sentence-level extraction and semantic alignment between source documents and summaries, thereby constructing traceable and verifiable decision evidence chains. The method introduces the first sentence-level alignment evaluation paradigm, overcoming the limitations of conventional summary-level scoring. It is the first approach to achieve state-of-the-art (SOTA) performance while providing transparent, step-by-step reasoning paths and demonstrating robustness against hallucination. On the SummEval benchmark, it attains a consistency score of 0.580—significantly outperforming the GPT-4 evaluator (0.521)—thereby advancing both evaluation accuracy and interpretability.

0 citationsRead paper
Recent publications

Latest Papers

AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume

Aug 29, 2025

This paper addresses core challenges in automatic text summarization evaluation—poor metric reproducibility, low correlation with human judgments, and the trade-off between computational cost and result stability. To this end, we introduce the first open-source, unified evaluation framework supporting standardized, reproducible comparison of diverse metrics, including ROUGE, G-Eval, SEval-Ex, and multiple LLM-based evaluators. Systematic experiments on SummEval reveal that LLM-based metrics exhibit substantial stochasticity and irreproducibility; quantitative analysis further confirms that high human alignment often comes at the cost of elevated computational overhead and reduced output stability. Our key contributions are: (1) the first unified framework compatible with heterogeneous evaluation paradigms; (2) the first empirical demonstration of reproducibility deficits in LLM-based evaluation; and (3) advancement of transparent, standardized evaluation protocols, establishing a reproducible benchmark for future summarization evaluation research.

0 citationsRead paper

SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation

May 04, 2025

Addressing the trade-off between performance and interpretability in summary evaluation, this paper proposes a fine-grained assessment framework based on atomic sentence decomposition: leveraging large language models (LLMs) to perform sentence-level extraction and semantic alignment between source documents and summaries, thereby constructing traceable and verifiable decision evidence chains. The method introduces the first sentence-level alignment evaluation paradigm, overcoming the limitations of conventional summary-level scoring. It is the first approach to achieve state-of-the-art (SOTA) performance while providing transparent, step-by-step reasoning paths and demonstrating robustness against hallucination. On the SummEval benchmark, it attains a consistency score of 0.580—significantly outperforming the GPT-4 evaluator (0.521)—thereby advancing both evaluation accuracy and interpretability.

0 citationsRead paper