Overview of the TREC 2025 Million Large Language Models track

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of expert selection within agent ecosystems, where massive numbers of large language models (LLMs) lack static metadata. To overcome this limitation, this work proposes an adaptive expert selection paradigm based on behavioral retrieval, shifting the retrieval target from documents to LLMs. Specifically, it dynamically infers model expertise by observing outputs and trains a ranking model leveraging queries, answers, and log-probability data to uncover latent capability representations. Furthermore, this research introduces the first large-scale benchmark and evaluation platform for AI expertise retrieval among agents. Experimental results validate the effectiveness of dynamic capability assessment, offering a novel solution for LLM routing in metadata-scarce scenarios.
📝 Abstract
Agentic AI envisions ecosystems of intelligent agents collaboratively solving complex tasks with minimal human intervention. In such ecosystems, each agent possesses specialized expertise, making effective expert selection central to overall system performance. While most current approaches assume a small number of well-documented models, real-world expertise is far more diverse and cannot be adequately captured through static metadata or hand-written descriptions. We anticipate a future with millions of specialized language models (LLMs), each excelling in different domains or problem types. Rather than relying on predefined capability statements, we propose a retrieval-based paradigm in which an assistant agent infers expertise dynamically by examining models'observable behavior. Upon receiving a user query, the assistant ranks candidate LLMs based on demonstrated competence, enabling efficient and adaptive expert selection. The TREC Million LLM Track operationalizes this paradigm by shifting the retrieval target from documents to expert LLMs. Participants are given a discovery set consisting of queries, answers, and log-probabilities from more than one thousand LLMs and are challenged to infer meaningful expertise representations for each model. Given an unseen test query, systems must then rank the LLMs according to their expected performance, providing the first large-scale benchmark for expertise retrieval in agentic AI.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Expert Selection
Large Language Models
Expertise Retrieval
Model Ranking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic AI
Expert Retrieval
Large Language Models
Retrieval-based Paradigm
Log-probabilities
🔎 Similar Papers
No similar papers found.