Token Counts Are Not Model Lineage: A Frozen-Threshold Holdout Study of Black-Box LLM API Fingerprinting

📅 2026-08-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过冻结阈值实验验证了token计数作为模型家族归属信号的有效性,发现其仅能作为共享分词栈的指纹,而非独立的必要测试。
📝 Abstract
Black-box model attribution is increasingly relevant when large language models (LLMs) are served through relay and reseller APIs. A tempting low-cost signal is the prompt-token count returned by an OpenAI-compatible endpoint: two models that share a tokenizer and chat template may produce the same count sequence up to a fixed offset. Yet the validity of this signal for broader \emph{model-family} attribution has received little direct holdout testing. We conduct a frozen-threshold study over 24 labeled endpoint pairs, split evenly into a development set and an untouched holdout set, with three temporal repeats and 30 controlled texts per pair. We introduce a validity-gated result contract that distinguishes an observed dissimilarity from an uninformative measurement caused by missing usage data, rate limits, or endpoint policy. The resulting shift-invariant exact-match score perfectly separates the 12 development pairs, yielding a frozen threshold of 0.725. On holdout, however, only 6 of 12 pairs are eligible under the pre-specified three-repeat rule. Among eligible pairs, balanced accuracy is 0.75, sensitivity is 0.50 (95\% Wilson interval 0.15--0.85), and specificity is 1.00 (0.342--1.00). Two same-family pairs---Qwen 3.8 and DeepSeek V4 variants---fall below the frozen threshold. Across 4,320 formal API calls, every log is replayable, while holdout contains 189 non-200 responses and 157 successful responses without prompt-token usage. The study therefore validates token-count consistency as a fingerprint of a shared \emph{tokenization stack}, but rejects its use as a standalone necessary test for model-family lineage.
Problem

Research questions and friction points this paper is trying to address.

Black-box model attribution
Token counts
Model lineage
LLM API Fingerprinting
Innovation

Methods, ideas, or system contributions that make the work stand out.

frozen-threshold holdout study
black-box LLM API fingerprinting
token-count consistency
validity-gated result contract
🔎 Similar Papers
2024-10-02arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
B
Bo Chen
Institute of Computing Technology, Chinese Academy of Sciences