When Harness Beats Scale, and When Reading Beats Both

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the imbalance between model scale and efficiency in document-based quantitative reasoning, alongside OCR degradation on test sets. We propose a reasoning pipeline integrating hybrid block retrieval, Program-of-Thought (PoT) generation, sandbox execution, and knowledge graph augmentation. Our findings reveal that smaller models trained for structured output can outperform larger counterparts, and that auditing input physical properties should take precedence over architectural design. Experimentally, a 27B-parameter model matches the performance of a 72B baseline while reducing carbon emissions by 75%. However, interference from watermarked PDFs causes a significant decline in test rankings, thereby validating the decisive impact of input quality on system robustness.
📝 Abstract
We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consistency sampling, and entity enrichment from chunk-level knowledge graphs. On our held-out split, application architecture moved the metrics far more than model scale did: PoT added 0.282 joint accuracy to a compact 7B model but at most 0.005 to a 72B model, and a 27B model with the full harness matched the 72B (0.884 vs.\ 0.873) at roughly 2.7$\times$ fewer parameters and a quarter of the CO$_2$. We read this through a distinction between world knowledge, which scales steeply with parameters, and language knowledge, which scales gently, and show that structured-output training makes a compact model harness-ready rather than merely small. On the raster, watermarked test PDFs the same system collapsed to 13.58\% joint (rank 149 of 163); a controlled re-rendering of the validation set reproduces the OCR half of the collapse while bounding what the simulation misses. Auditing the physical nature of evaluation inputs precedes architecture, and the leaderboard's bimodality is consistent with reading quality, not reasoning, having separated the field.
Problem

Research questions and friction points this paper is trying to address.

document-grounded quantitative reasoning
model scale vs architecture
OCR quality
Program-of-Thoughts
evaluation robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Program-of-Thoughts
Hybrid Block Retrieval
Knowledge Graph
Structured-Output Training
OCR Audit
🔎 Similar Papers
No similar papers found.