LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the capacity of general-purpose large language models (LLMs) to directly output calibrated decision probabilities and the necessity of fine-tuning. We propose LLM2Jev, a framework that extracts next-token probabilities for architecture-preserving decision elicitation, supporting both training-free inference and KL-anchored LoRA fine-tuning. Our findings demonstrate that strong base models achieve zero-shot performance comparable to specialized decision models, with fine-tuning gains diminishing as model capability increases. Furthermore, integrating a tree-decomposed listwise loss effectively mitigates behavioral degradation. Experiments reveal that a 4B-parameter model in the zero-shot setting matches community benchmarks and outperforms letter-logit approaches, while LoRA substantially enhances weaker models without compromising generation quality.
📝 Abstract
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Decision Models
Fine-tuning
Probability Calibration
Categorical Distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decision Models
Next-token Probabilities
Listwise Loss
KL Divergence Anchoring
Fine-tuning
💼 Related Jobs
No related jobs found.
Yinheng Li
Yinheng Li
Microsoft
J
Justin Wagle
Microsoft