SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost and low efficiency of LLM-mediated retrieval within large-scale agent skill libraries by proposing a two-stage skill retrieval framework grounded in standard information retrieval (IR) pipelines. The proposed method integrates a BGE-base bi-encoder, a cross-encoder, and the BM25 algorithm, with its interfaces exposed via the Model Context Protocol (MCP). Experimental results demonstrate that conventional IR techniques achieve performance comparable to LLM-mediated loops on the SkillsBench benchmark while substantially reducing token consumption. Specifically, this approach decreases the per-task cost from $51.30 to $27.54, offering a cost-effective and highly efficient alternative for agent skill retrieval.
📝 Abstract
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
Problem

Research questions and friction points this paper is trying to address.

Agent Skill Retrieval
Large Language Model
Information Retrieval
Cost Reduction
Skill Marketplace
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Skill Retrieval
Two-stage Retriever
Cross-encoder
Information Retrieval
Cost Efficiency