Similarity-Distance-Magnitude Language Models

📅 2025-10-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the low statistical efficiency of large language models (LLMs) in instruction-following tasks, caused by frequent abstention—i.e., refusal to generate outputs despite being capable. To tackle this, we propose the Similarity–Distance–Magnitude (SDM) language model. Methodologically, we adapt a pretrained decoder-only Transformer via supervised fine-tuning into a sequence prediction model, introducing an SDM activation function at the output layer to explicitly partition high-probability regions into trustworthy generation intervals. We further enhance calibration and decision boundary sharpness through contrastive input encoding and online hard negative sampling, coupled with an SDM-based basis-transformed loss function. Our core contribution is the first integration of the SDM architecture into the final decoder layer, enabling explicit modeling of generation confidence. Experiments demonstrate substantial reductions in abstention rates, outperforming strong supervised baselines across multiple instruction-following benchmarks while improving both generation effectiveness and probabilistic calibration.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsReasoning under Uncertainty: Relational Probabilistic Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
We introduce Similarity-Distance-Magnitude (SDM) language models (LMs), which are sequence prediction models fine-tuned to maximize the proportion of generations in the well-calibrated, high-probability region partitioned by a final-layer SDM activation layer used for binary classification of instruction-following. We demonstrate that existing pre-trained decoder-only Transformer LMs can be readily converted into SDM LMs via supervised fine-tuning, using the final-layer SDM activation layer during training to estimate a change-of-base for a supervised next-token loss over a contrastive input encoding scheme, with additional hard negative examples generated online during training. This results in reduced abstentions (i.e., improved statistical efficiency) compared to strong supervised baselines.
Problem

Research questions and friction points this paper is trying to address.

Fine-tune models to maximize calibrated high-probability generations
Convert pre-trained Transformers using contrastive input encoding
Reduce abstentions and improve statistical efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

SDM activation layer enables binary classification
Fine-tunes Transformers with contrastive input encoding
Online hard negative generation reduces abstentions
🔎 Similar Papers