HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the privacy-cost-accuracy trade-off inherent in locally deployed small language models. We propose a novel hidden-state-based self-checking mechanism, wherein fine-tuning endows intermediate representations with predictive capabilities regarding answer correctness. Building upon this, we design a cascaded inference strategy that leverages prefilling and state-guided routing to dynamically allocate computational resources and human intervention, thereby enabling efficient and secure decision-making without requiring cloud-based upgrades. Empirical evaluations demonstrate that on single-device deployments, our approach reduces latency by 2.7× while achieving accuracy comparable to chain-of-thought reasoning. Furthermore, in multi-model server scenarios, it outperforms mainstream cascading methods, yielding substantially lower error rates.
📝 Abstract
Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model's own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.
Problem

Research questions and friction points this paper is trying to address.

local language model deployment
inference efficiency
safety
query deferral
hidden states
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inference-Time Self-Checks
Hidden States
Local Language Model Deployment
Cascade Policy
Efficiency and Safety
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.