Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

📅 2023-06-01
🏛️ arXiv.org
📈 Citations: 18
Influential: 0
📄 PDF

career value

191K/year
🤖 AI Summary
Current large language model (LLM) evaluations suffer from methodological homogeneity and detachment from real-world societal needs. Method: This paper reframes LLM evaluation around the core objective of *socio-technical gap*—systematically measuring an LLM’s capacity to fulfill diverse, authentic application requirements. We introduce the first interdisciplinary evaluation paradigm integrating human–computer interaction (HCI), explainable AI (XAI), and natural language generation (NLG), structured along three dimensions: context-driven scenario design, stakeholder需求 mapping, and feasibility-aware trade-off analysis. Contribution/Results: We propose the first conceptual framework for the socio-technical gap, identifying critical deficiencies in existing benchmarks; establish a principled, real-world–oriented evaluation pathway; and articulate foundational open research questions. By shifting focus from technical metrics to human-centered outcomes, this paradigm advances LLM assessment toward equitable, socially grounded deployment and actively supports narrowing the socio-technical gap.
📝 Abstract
The recent development of generative large language models (LLMs) poses new challenges for model evaluation that the research community and industry have been grappling with. While the versatile capabilities of these models ignite much excitement, they also inevitably make a leap toward homogenization: powering a wide range of applications with a single, often referred to as ``general-purpose'', model. In this position paper, we argue that model evaluation practices must take on a critical task to cope with the challenges and responsibilities brought by this homogenization: providing valid assessments for whether and how much human needs in diverse downstream use cases can be satisfied by the given model ( extit{socio-technical gap}). By drawing on lessons about improving research realism from the social sciences, human-computer interaction (HCI), and the interdisciplinary field of explainable AI (XAI), we urge the community to develop evaluation methods based on real-world contexts and human requirements, and embrace diverse evaluation methods with an acknowledgment of trade-offs between realisms and pragmatic costs to conduct the evaluation. By mapping HCI and current NLG evaluation methods, we identify opportunities for evaluation methods for LLMs to narrow the socio-technical gap and pose open questions.
Problem

Research questions and friction points this paper is trying to address.

Generative Language Models
Socio-technical Gap
Real-world Application
Innovation

Methods, ideas, or system contributions that make the work stand out.

Socio-technical Gap Assessment
Interdisciplinary Evaluation Framework
Cost-effective Evaluation Methods