🤖 AI Summary
This study addresses a critical decision bottleneck in the Jev model for arithmetic tasks, wherein the model exhibits verification capabilities yet fails to correctly reject erroneous options, particularly when no valid answer is present. By identifying the discrepancy between native Boolean verification and final selection, this work proposes a lightweight calibration strategy that optimizes the decision threshold using a development set, thereby resolving the bottleneck without requiring retraining. The proposed approach substantially improves arithmetic rejection accuracy from 7% to 79% while maintaining a 97% accuracy rate when correct answers are available. Consequently, this method effectively mitigates model failure without introducing additional inference overhead.
📝 Abstract
Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefined options. When candidate sets contain no valid answer, TypeSafe recommends including an "other" or "none-of-the-above" option to enable rejection. In this report, however, we identify an arithmetic-dependent rejection bottleneck: Jev reliably selects correct numerical answers when available but frequently accepts incorrect alternatives when they are absent despite an explicit rejection option. On paired arithmetic problems, answer-present accuracy reaches 99%, while correct rejection falls to 7%. Moreover, this gap persists across numerical magnitudes, operation depths, contextual formulations, and rejection labels, and extends to scenarios such as time calculation and capacity rounding. Yet native Boolean verification achieves 99% exact-match accuracy on the same answer-absent arithmetic cases, showing that categorical rejection can fail even when the model successfully verifies candidate correctness. Finally, we show that a simple decision threshold selected on separate development problems raises arithmetic rejection accuracy from 7% to 79% while retaining 97% answer-present accuracy, substantially mitigating the failure without retraining or additional inference.