Score
Designs and implements fine-tuning procedures for encoder or siamese architectures that use contrastive or energy-based losses to produce and align vector embeddings from paired or mined positives/negatives, optimizing them for semantic similarity, dense retrieval, and reranking. Builds supervised and augmentation-based pair generation and negative-sampling pipelines, tunes loss/temperature/energy parameters, and evaluates model behavior with retrieval and similarity metrics (e.g., recall@k) and robustness to augmentation or language shifts.
This work addresses the high computational cost of conventional neural architecture search (NAS) performance predictors, which often rely on expensive fine-tuning or intricate architecture representations. The authors propose Code-Oriented Language Model Embeddings (COLE), a method that directly uses raw PyTorch class definition code as input and leverages a frozen off-the-shelf language model to extract architecture embeddings. Coupled with a lightweight regression head, COLE constructs an efficient performance predictor without requiring NAS-specific fine-tuning. Evaluated on the NAS-Bench-201 benchmark, COLE achieves within 1% of the optimal architecture’s accuracy while reducing the evaluation budget by 34% compared to path-based encoding. Furthermore, experiments on CIFAR-100 demonstrate its strong generalization capability and superior search efficiency.
This study systematically compares contrastive loss and triplet loss in audio-visual cross-modal embedding learning, focusing on their representational capacity differences. To characterize their intrinsic distinctions—particularly in intra-class variance control, hard sample mining, and optimization dynamics—we propose a quantitative analysis framework measuring loss decay rate, positive-pair activation ratio, and gradient norm. Controlled experiments across MNIST, CIFAR-10, CUB-200, CARS196, and synthetic datasets demonstrate that triplet loss prioritizes hard samples, preserves richer semantic details, and significantly improves fine-grained classification and cross-modal retrieval performance; contrastive loss yields more compact yet less discriminative embeddings. Crucially, this work provides the first optimization-dynamics–driven mechanistic explanation of how these losses differentially shape representation quality—establishing both theoretical grounding and empirical evidence for principled loss selection in metric learning.
Self-supervised learning faces two key challenges: Siamese networks suffer from representational collapse, while contrastive learning relies on negative samples, leading to poor robustness under small batch sizes. To address these issues, this paper proposes a negative-sample-free implicit contrastive learning paradigm. Its core innovation is a guided stop-gradient mechanism that dynamically blocks gradients between symmetric positive sample pairs, thereby implicitly encoding contrastive signals without requiring negative samples, prediction heads, or asymmetric encoders. The method is fully compatible with SimSiam and BYOL frameworks, needing only standard momentum updates and symmetric loss functions. Extensive experiments demonstrate significant performance gains on ImageNet, robust training with extremely small batch sizes (e.g., 8), and superior training stability and generalization compared to existing negative-sample-free approaches.
Vision-language models like CLIP align image and text embeddings but lack semantic comparability and analogical reasoning capabilities in their embedding space, hindering vectorized reasoning about inter-image differences. Method: We propose a contrastive learning fine-tuning framework that explicitly aligns image embedding differences to LLM-generated textual difference descriptors (e.g., “thinner”, “brighter”), enabling native pairwise image difference reasoning within CLIP’s embedding space. We further introduce a novel “comparative prompting” inference paradigm to restructure the embedding geometry. Contribution/Results: After fine-tuning on synthetic difference data, our method improves average accuracy by 3.2% across attribute ranking, zero-shot classification, and retrieval tasks. It significantly enhances linear analogy preservation and directional consistency in the embedding space—providing stronger geometric foundations for downstream applications such as text-to-image generation.
To address poor generalization, parameter redundancy, and weak cross-task transferability in fine-tuned models, this paper proposes a neural parameter learnable pruning method grounded in the low-rank subspace spanned by task vectors. Unlike conventional structured pruning, our approach jointly optimizes critical parameter masks within this task-vector-derived low-rank subspace—simultaneously suppressing catastrophic forgetting, ensuring model interpolation compatibility, and maximizing compression efficiency. By integrating task vector modeling, subspace projection, and differentiable mask search, we achieve synergistic lightweight pruning and cross-domain knowledge transfer. Extensive experiments on vision, NLP, and multimodal benchmarks demonstrate that our method retains near-original accuracy under high compression ratios (>50%), significantly outperforming state-of-the-art pruning and model merging techniques. The implementation is publicly available.
This work addresses the "bag-of-concepts" effect in dual-encoder vision-language models, which arises from similarity aggregation mechanisms and undermines compositional reasoning for logically constrained queries (e.g., “umbrella and no person”). To overcome this limitation, the authors propose a factorized reasoning framework that decouples concept evidence extraction from logical constraint enforcement. They introduce Logic Constraint Score Editing (LCSE), a training-free method that explicitly performs logical inference on top of frozen encoders. The study identifies similarity aggregation as the primary cause of compositional failure and introduces FACTOR-Bench, a new evaluation benchmark. Experiments show that LCSE achieves 85.5% accuracy on FACTOR-Bench (90.7% with SigLIP-2), substantially outperforming the best fine-tuned baseline at 73.2%, and improves accuracy from 27.2% to 65.2% on NegBench COCO MCQ while preserving original retrieval performance.
This work investigates whether the standard Transformer architecture is universally optimal across all tasks and proposes an architectural refinement that introduces task-specific inductive biases through learnable nonlinear components, such as GeLU or softmax. The approach preserves the original Transformer structure while replacing key activation functions with task-optimized counterparts learned during training. Experimental results demonstrate that this modification substantially improves learning speed, in- and out-of-distribution generalization, and training stability on algorithmic reasoning tasks. Consistent, albeit more modest, performance gains are also observed in language and code modeling, accompanied by enhanced cross-domain transfer capabilities. These findings indicate that the standard Transformer is not locally optimal for specific tasks and that incorporating task-tailored design elements—despite a trade-off in generality—can effectively boost performance.
This work addresses the limitation of conventional cross-entropy training under teacher forcing, which optimizes only token-level predictions and fails to align with sequence-level behavior in autoregressive generation. To overcome this, the authors propose Energy-Based Fine-Tuning (EBFT), a novel approach that integrates energy-based modeling with semantic feature matching. EBFT generates multiple candidate sequences in parallel, extracts batch-wise feature embeddings, and performs online updates via policy gradient optimization, augmented with KL regularization to directly shape sequence-level statistical properties. Notably, it provides dense semantic feedback without requiring task-specific verifiers or preference models. Experiments across question answering, unstructured code generation, and machine translation demonstrate that EBFT achieves higher downstream accuracy than supervised fine-tuning (SFT), matches the performance of RLVR, and yields lower validation cross-entropy.
This study investigates how to select optimal loss function–optimizer pairings for robust training across structurally diverse neural networks. Leveraging the LEMUR heterogeneous architecture pool, the authors systematically evaluate 18 combinations of three loss functions—Cross-Entropy, Negative Log-Likelihood (NLL), and Normalized Gradient Loss (NGL)—with six optimizer families, including SGD and variants of Adam, across six image classification benchmarks. By fixing hyperparameters to control for confounding variables, they generate 594 model variants. Their large-scale empirical analysis reveals, for the first time in heterogeneous architectures, that loss–optimizer pairings significantly impact performance, challenging the notion of a universally effective combination: Cross-Entropy paired with Adam or AdamW demonstrates the most consistent robustness; NGL is only effective with adaptive optimizers and rivals Cross-Entropy in convolutional models; whereas Adagrad and Adadelta consistently underperform.
Although supervised fine-tuning (SFT) exerts only a subtle effect on the cosine similarity of hidden activations in large language models, their internal representations may nonetheless undergo substantial changes. This work proposes a high-resolution mechanistic analysis framework based on pretrained sparse autoencoders (SAEs), integrating representational geometry with layer-wise feature tracking. For the first time, it reveals systematic semantic feature shifts induced by SFT within sparse latent spaces and identifies layer-update patterns uniquely associated with safety alignment. The method precisely localizes key semantic features whose distributions are altered by SFT. All code and analyses are publicly released.