What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

📅 2026-09-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
"This study addresses the presence of bias in the unpatched probability samples within the MCP server ecosystem. It randomly selected 400 npm/stdio servers, probed using publicly available seeds, to evaluate their initialization handshake success rates, JSON Schema compliance, and tool description redundancy. For the first time, it reveals the true state of unpatched samples, identifying primary issues such as server startup failures and significant discrepancies in tool descriptions across benchmark sets. The methodology included random sampling, network probing, JSON Schema validation, and cosine similarity. The results indicate that only 48.8% of the servers successfully completed the initialization handshake; among the 195 operational servers, the tool-level omission rate was 58.8%; and the cross-author approximate duplication rate for genuine MCP tools was 0%, while for BFCL v4, it was 16.7%."
📝 Abstract
Studies of the Model Context Protocol (MCP) server ecosystem draw their samples in ways that quietly select for servers that work: reference sets, popularity lists, hand-curated frames, or pipelines that repair a server until it starts. We report what an unrepaired probability sample actually contains. From a 24,135-server registry census we draw 400 npm/stdio servers with a published seed and probe each one over the wire. Only 48.8% complete an initialize handshake, against 66.7% for a hand-curated frame measured with the same instrument, and the dominant failure is not missing credentials (13.3%) but servers that never start at all (37.5%). Among the 195 that do run, hard conformance is total: zero fatal JSON Schema violations across 2,766 advertised tools. Optional safety annotations are the real variance, and the tool-level omission rate on a random draw is 58.8% against 41.5% on the curated frame, so curation flatters this figure too. We then compare the tool descriptions these servers advertise against two tool-use benchmark corpora using one method held constant. Real MCP tools show 2.8% near-duplication at cosine 0.70, and all of it lies within single servers: cross-author near-duplication is 0.0% at every threshold tested. BFCL v4 shows 16.7%, of which 16.4 points lie between independently presented tasks. UltraTool shows 0.3%, cleaner than real tools, so this is a property of BFCL and not of synthetic corpora as a class. Separately, 68.8% of raw BFCL rows and 85.6% of raw UltraTool rows are exact name-plus-description repeats, against 0.4% for real MCP, so any statistic computed over these releases without global deduplication measures repetition rather than tools. All figures regenerate from released scripts and a published seed.
Problem

Research questions and friction points this paper is trying to address.

MCP Registry
Random Draw
Tool-Use Benchmarks
Server Ecosystem
Conformance
Innovation

Methods, ideas, or system contributions that make the work stand out.

unrepaired probability sample
initialize handshake
tool-level omission rate
near-duplication
global deduplication
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haseeb Mohammed Afsar
Independent researcher