Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent failure of small open-source models in local agent frameworks, where they frequently struggle with real-world tasks due to context overflow, divergent self-correction, and tool-calling loops. To overcome these limitations, this work proposes a local-first agent framework targeting Windows and Ollama environments. It introduces ten compensatory mechanisms—including byte-level net-zero prefill budgeting, task-rereading completion gating, and signature-level loop detection—integrated with refined context management and deterministic artifact scoring to systematically mitigate small-model failure modes. Evaluated on the LRAB benchmark, the proposed framework achieves a score of 0.886, significantly outperforming mainstream alternatives. Furthermore, it attains 0.856 on τ²-bench, demonstrating its effectiveness in substantially enhancing the reliability and task completion rates of small language models executing real-world tasks.
📝 Abstract
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses $\times$ 4 open models (2B-35B) $\times$ 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on $τ^2$-bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.
Problem

Research questions and friction points this paper is trying to address.

small open models
agent harness
local-first
task completion
context overflow
Innovation

Methods, ideas, or system contributions that make the work stand out.

Local-first agent harness
Small open models
Net-zero prefill budget
Finish gate
Signature-level loop detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hao Wang
School of Mathematics and Physics, University of Science and Technology Beijing
T
Ting Huang
Honor Device Co., Ltd.