Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the overestimation of small models’ tool-calling capabilities caused by keyword-matching dependencies in existing benchmarks. We propose a low-cost diagnostic cascade integrating first-token probing, embedding drift detection, and factor analysis, revealing that specific training stages erase tool-calling priors. Targeted supervised fine-tuning (SFT) is subsequently applied for precise remediation. Experiments demonstrate that the model’s effective invocation rate increases from 0.1 to 0.959, while zero-shot pass rates significantly outperform baselines, confirming that the intervention preserves underlying representations. This work provides a systematic framework for evaluating and restoring the genuine tool-calling abilities of small language models under limited computational budgets.
πŸ“ Abstract
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ on <|tool_call|>), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 ($p = 0.004$). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.
Problem

Research questions and friction points this paper is trying to address.

small language models
tool use
keyword-matching benchmarks
false positives
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tool-use evaluation
Diagnostic ladder
Targeted SFT
First-token probe
Small language models
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
J
Juan S. Santillana
Independent Researcher