From Understanding to Generation: An Efficient Shortcut for Evaluating Language Models

📅 2025-06-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Evaluating natural language generation (NLG) capabilities—such as mathematical reasoning, code generation, factual knowledge, and reading comprehension—during large language model (LLM) training incurs high computational cost and impedes frequent monitoring. Method: This work systematically reformulates expensive NLG evaluation tasks into lightweight natural language understanding (NLU) multiple-choice formats. We propose a general task rephrasing paradigm and a cross-format correlation analysis framework to enable efficient capability assessment. Contribution/Results: We empirically validate, across eight diverse LLMs and four capability dimensions, a strong correlation (average Spearman ρ > 0.92) between NLG and NLU evaluations, establishing their functional interchangeability. Our approach achieves an average 35.2× speedup in evaluation latency, drastically reducing training-time monitoring overhead while enabling scalable, dynamic LLM assessment.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Iterative evaluation of LLMs during training is essential to ensure expected capability development, but can be time- and compute-intensive. While NLU tasks, where the model selects from fixed answer choices, are cheap to evaluate, essential capabilities like reasoning and code generation rely on the more time-consuming NLG (token-by-token generation) format. In this work, our aim is to decrease the computational burden of NLG benchmarks in order to enable monitoring crucial LLM capabilities during model training. We reformulate generative tasks into computationally cheaper NLU alternatives. We test the performance correlation between the original and reformulated tasks using 8 LMs of various sizes and 4 capabilities: mathematical reasoning, code generation, factual knowledge and reading comprehension. Our results show a strong correlation between task formats, supporting capability assessment via cheaper alternatives and achieving over 35x average reduction in evaluation time. We plan to publish our benchmark adaptions.
Problem

Research questions and friction points this paper is trying to address.

Reducing computational cost of NLG benchmarks for LLM training
Converting generative tasks into cheaper NLU alternatives
Validating performance correlation between original and reformulated tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reformulate generative tasks into NLU
Test performance correlation between formats
Achieve 35x evaluation time reduction