🤖 AI Summary
Although general-purpose large language models (LLMs) exhibit broad capabilities, the relative advantages of specialized decision-making models in knowledge-intensive tasks remain unclear. This study addresses this gap by conducting a large-scale comparative evaluation of Jev, a specialized decision-making model grounded in the System One paradigm, against 19 mainstream LLMs using multiple-choice benchmarks. Results indicate that Jev achieves top-tier performance on knowledge-centric tasks such as MMLU-Redux and ARC-Challenge, rivaling frontier models. However, it significantly underperforms the median of frontier models on multi-step mathematical reasoning tasks like MathQA. These findings demonstrate that specialized decision-making models can effectively substitute general-purpose LLMs in knowledge-dependent scenarios while delineating their limitations in complex computational reasoning.
📝 Abstract
Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.