🤖 AI Summary
This study investigates the capacity of general-purpose large language models (LLMs) to directly output calibrated decision probabilities and the necessity of fine-tuning. We propose LLM2Jev, a framework that extracts next-token probabilities for architecture-preserving decision elicitation, supporting both training-free inference and KL-anchored LoRA fine-tuning. Our findings demonstrate that strong base models achieve zero-shot performance comparable to specialized decision models, with fine-tuning gains diminishing as model capability increases. Furthermore, integrating a tree-decomposed listwise loss effectively mitigates behavioral degradation. Experiments reveal that a 4B-parameter model in the zero-shot setting matches community benchmarks and outperforms letter-logit approaches, while LoRA substantially enhances weaker models without compromising generation quality.
📝 Abstract
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.