A Practical Guide for Evaluating LLMs and LLM-Reliant Systems

📅 2025-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing synthetic benchmarks and static metrics inadequately assess LLM-dependent systems in real-world settings. To address this, we propose an end-to-end evaluation framework driven by user needs: it integrates task-scenario analysis and human-AI collaborative annotation to construct representative datasets; defines multidimensional utility metrics—practicality, robustness, and explainability—and embeds them into the development-deployment feedback loop to enable dynamic iteration and A/B-test-driven empirical evaluation. This work is the first to systematically bridge engineering practice with evaluation science, transcending limitations of conventional paradigms. Validated across multiple industrial-scale LLM applications, our framework significantly improves the correlation between evaluation outcomes and both user satisfaction and core business metrics, thereby robustly supporting the reliable productization of LLM systems.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Humans and AI: User Experience and UsabilityPhilosophy and Ethics of AI: Safety, Robustness & Trustworthiness

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating success
📝 Abstract
Recent advances in generative AI have led to remarkable interest in using systems that rely on large language models (LLMs) for practical applications. However, meaningful evaluation of these systems in real-world scenarios comes with a distinct set of challenges, which are not well-addressed by synthetic benchmarks and de-facto metrics that are often seen in the literature. We present a practical evaluation framework which outlines how to proactively curate representative datasets, select meaningful evaluation metrics, and employ meaningful evaluation methodologies that integrate well with practical development and deployment of LLM-reliant systems that must adhere to real-world requirements and meet user-facing needs.
Problem

Research questions and friction points this paper is trying to address.

Evaluate LLM systems in real-world scenarios
Address challenges beyond synthetic benchmarks
Develop practical evaluation frameworks for deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proactive curation of representative datasets
Selection of meaningful evaluation metrics
Integration with practical development methodologies