Comparative Approaches to Agent Retrieval over Large Skill Libraries

📅 2026-08-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficient skill selection and ranking from large skill repositories, where full loading incurs high costs and existing systems lack autonomous scheduling mechanisms. The authors compare two retrieval approaches: a hybrid ranker combining lexical and dense embeddings that enables on-demand sparse loading, and a typed workflow knowledge graph encoding preconditions and dataflow dependencies. Experimental results show that augmenting the strong ranker with graph structure does not improve retrieval effectiveness—graph candidates are confined to the embedding neighborhood already covered by the ranker, limiting scope expansion. On 117 real-world queries, the hybrid ranker achieves a Top-5 hit rate of 73.5% ± 8.0%, whereas the knowledge graph underperforms by 11.2 points (p = 0.0007) under the same token budget and fails to recover 73% of queries missed by the ranker. The study also reveals that self-authored queries can overestimate hit rates by up to 44 percentage points, exposing significant evaluation bias.
📝 Abstract
Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and provides no structure for autonomous sequencing. We study two systems for this problem over a corpus of 690 skills: a hybrid ranker combining lexical and dense-embedding retrieval for sparse, on-demand loading, and a typed knowledge graph encoding workflow relations such as prerequisites, data flow, and ordering. On a set of 117 realistic, non-echoing queries, the hybrid ranker retrieves the correct skill within the top five in 73.5% +/- 8.0 of cases, leaving roughly a quarter of queries unserved. When used as the design intended (substituting graph neighbours for additional ranked results at matched token budget), the graph is significantly worse (-11.2 points, p = 0.0007). Its LLM-generated edge layer adds nothing over neighbours obtained free from a local embedding pass, and 73% of the queries the ranker misses are not reachable through the graph at all. We attribute this to a pre-filter topology bound. Because the graph's candidate edges are drawn from the same embedding neighbourhood the ranker already searches, 98.6% of typed edges connect skills the ranker had already surfaced together. The graph can enrich relation semantics but cannot extend retrieval reach. We further show that evaluating on author-written queries overstates hit@5 by up to 44 points, which would have hidden these results entirely. Our contribution is a mechanistic account of why added structure does not improve retrieval over a strong ranker, and identify the conditions under which adding structural interdependence into the retrieval is optimal.
Problem

Research questions and friction points this paper is trying to address.

agent retrieval
skill libraries
retrieval ranking
knowledge graph
sparse loading
Innovation

Methods, ideas, or system contributions that make the work stand out.

agent retrieval
skill library
hybrid ranker
typed knowledge graph
retrieval evaluation bias