🤖 AI Summary
This work addresses the low statistical efficiency of large language models (LLMs) in instruction-following tasks, caused by frequent abstention—i.e., refusal to generate outputs despite being capable. To tackle this, we propose the Similarity–Distance–Magnitude (SDM) language model. Methodologically, we adapt a pretrained decoder-only Transformer via supervised fine-tuning into a sequence prediction model, introducing an SDM activation function at the output layer to explicitly partition high-probability regions into trustworthy generation intervals. We further enhance calibration and decision boundary sharpness through contrastive input encoding and online hard negative sampling, coupled with an SDM-based basis-transformed loss function. Our core contribution is the first integration of the SDM architecture into the final decoder layer, enabling explicit modeling of generation confidence. Experiments demonstrate substantial reductions in abstention rates, outperforming strong supervised baselines across multiple instruction-following benchmarks while improving both generation effectiveness and probabilistic calibration.
📝 Abstract
We introduce Similarity-Distance-Magnitude (SDM) language models (LMs), which are sequence prediction models fine-tuned to maximize the proportion of generations in the well-calibrated, high-probability region partitioned by a final-layer SDM activation layer used for binary classification of instruction-following. We demonstrate that existing pre-trained decoder-only Transformer LMs can be readily converted into SDM LMs via supervised fine-tuning, using the final-layer SDM activation layer during training to estimate a change-of-base for a supervised next-token loss over a contrastive input encoding scheme, with additional hard negative examples generated online during training. This results in reduced abstentions (i.e., improved statistical efficiency) compared to strong supervised baselines.