🤖 AI Summary
Existing synthetic benchmarks and static metrics inadequately assess LLM-dependent systems in real-world settings. To address this, we propose an end-to-end evaluation framework driven by user needs: it integrates task-scenario analysis and human-AI collaborative annotation to construct representative datasets; defines multidimensional utility metrics—practicality, robustness, and explainability—and embeds them into the development-deployment feedback loop to enable dynamic iteration and A/B-test-driven empirical evaluation. This work is the first to systematically bridge engineering practice with evaluation science, transcending limitations of conventional paradigms. Validated across multiple industrial-scale LLM applications, our framework significantly improves the correlation between evaluation outcomes and both user satisfaction and core business metrics, thereby robustly supporting the reliable productization of LLM systems.
📝 Abstract
Recent advances in generative AI have led to remarkable interest in using systems that rely on large language models (LLMs) for practical applications. However, meaningful evaluation of these systems in real-world scenarios comes with a distinct set of challenges, which are not well-addressed by synthetic benchmarks and de-facto metrics that are often seen in the literature. We present a practical evaluation framework which outlines how to proactively curate representative datasets, select meaningful evaluation metrics, and employ meaningful evaluation methodologies that integrate well with practical development and deployment of LLM-reliant systems that must adhere to real-world requirements and meet user-facing needs.