General Decision Models: Benchmarking and Insights Beyond Jev

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the reliability limitations and deficient uncertainty estimation of generalist decision-making models in complex, long-horizon interactions. We introduce JEVal, the first bilingual benchmark for systematic evaluation across 25 models, and propose Reasoning-to-Readout self-distillation, a technique that internalizes multi-step reasoning from large language models into single-pass first-token decisions to yield the efficient InnerJev architecture. Experimental results demonstrate that InnerJev-27B achieves performance comparable to Jev while reducing typical query latency to merely 0.1 seconds. Furthermore, our quantitative analysis reveals that localized rapid decision-making systematically degrades success rates in long-horizon tasks, providing critical empirical evidence for bias analysis in complex system simulations.
πŸ“ Abstract
General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVal, a bilingual benchmark comprising 11,257 instances from 36 datasets across 10 application domains, and evaluate 25 model configurations spanning general decision models and generative LLMs. Our results show that (1) general decision models are most competitive when decisions can be resolved from available evidence, but weaken when they require specialist knowledge or faithful uncertainty estimation: they can often identify the most likely outcome while substantially overstating its probability. (2) In more dynamic and realistic systems involving long-horizon, multi-step interactions, the advantages of fast local decision making are offset by reliability failures at the system level. on $\tau$-bench, faster local decisions reduce median episode time but lower task success as decision errors accumulate over long trajectories. (3) In large-scale social simulation, decision models approach strong generative LLMs on individual response prediction at substantially lower inference cost, yet remain weaker in user profiling and exhibit larger aggregate estimation errors and systematic bias. Finally, we propose InnerJev-4B and InnerJev-27B, which internalize an open-weight LLM's own reasoning into a single-pass first-token decision through Reasoning-to-Readout Self-Distillation, with InnerJev-27B performing on par with Jev on JEVal while answering a typical query in about 0.1 s.
Problem

Research questions and friction points this paper is trying to address.

General Decision Models
Benchmarking
System-level Reliability
Social Simulation
Uncertainty Estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

General Decision Models
Reasoning-to-Readout Self-Distillation
JEVal Benchmark
InnerJev
First-token Decision
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Feiyu Duan
Feiyu Duan
Beihang University
natural language processing
Jiayu Lin
Jiayu Lin
Fudan University
J
Jia Wang
Shanghai Innovation Institute, Tongji University
J
Jun Xiang
Fudan University
Jialiang Wu
Jialiang Wu
Unknown affiliation
Xinnong Zhang
Xinnong Zhang
Fudan University
Natural Language ProcessingComputational Social ScienceLarge Language Models
H
Hanqi Yan
King’s College London
S
Siyuan Wang
Chinese University of Hong Kong
Z
Zhongyu Wei
Shanghai Innovation Institute, Fudan University