Similarity-Distance-Magnitude Language Models
This work addresses the low statistical efficiency of large language models (LLMs) in instruction-following tasks, caused by frequent abstention—i.e., refusal to generate outputs despite being capable. To tackle this, we propose the Similarity–Distance–Magnitude (SDM) language model. Methodologically, we adapt a pretrained decoder-only Transformer via supervised fine-tuning into a sequence prediction model, introducing an SDM activation function at the output layer to explicitly partition high-probability regions into trustworthy generation intervals. We further enhance calibration and decision boundary sharpness through contrastive input encoding and online hard negative sampling, coupled with an SDM-based basis-transformed loss function. Our core contribution is the first integration of the SDM architecture into the final decoder layer, enabling explicit modeling of generation confidence. Experiments demonstrate substantial reductions in abstention rates, outperforming strong supervised baselines across multiple instruction-following benchmarks while improving both generation effectiveness and probabilistic calibration.