Score
Design and implement representation‑learning methods that create and maintain prototypes or semantic anchors (knowledge‑guided, latent, spatial, local, and ordinal) and align model features to those prototypes using contrastive losses or regularizers. This competence covers building offline anchors, encoding ordinal or spatial structure in prototypes, and integrating external knowledge to guide prototype construction and feature alignment (e.g., via LoPA/local prototype alignment or prototype contrastive learning).
In federated learning, dual heterogeneity—arising from both non-IID data and divergent local models—induces a vicious cycle of representation inconsistency, classifier divergence, and prototype misalignment, hindering effective prototype sharing. To address this, we propose a semantic-anchor-driven federated prototypical learning framework. Our core innovation replaces locally biased prototypes with cross-client semantically consistent semantic anchors, decoupling prototype generation from representation learning. We further introduce anchor regularization and classifier calibration to jointly enforce intra-class compactness, inter-class separability, and decision boundary alignment. Additionally, we incorporate margin-augmented contrastive learning and iterative unified representation optimization. Evaluated on multiple image classification benchmarks, our method consistently outperforms existing prototypical federated learning approaches, achieving absolute accuracy gains of 3.2–7.8%.
This study addresses the insufficient discriminability of general-purpose text embeddings in principled evaluation caused by semantic overlap. To mitigate this, we propose Prototype-Guided Contrastive Learning (PGCL), which integrates prototype-anchored attention with shifted margin regularization to generate compact, task-adaptive representations while keeping the encoder frozen. This approach effectively resolves confusion between semantically similar samples with distinct task labels. Experimental results demonstrate that PGCL outperforms original frozen embeddings across three benchmark datasets, achieving particularly significant improvements on Amazon Reviews and matching strong baselines on other tasks. These findings validate the effectiveness of PGCL for parameter-efficient adaptation, offering a robust solution for enhancing embedding distinctiveness without extensive model retraining.
This study addresses the limitation of fixed geometric structures in contrastive learning, which struggle to accommodate context-dependent semantic similarity. To overcome this, we propose anchor divergence, a method that integrates contrastive learning, exponential family theory, and information geometry. By establishing a correspondence between anchor distributions and Bregman geometry, our approach defines a context-aware dynamic semantic geometry over fixed representations. Specifically, it directly controls the geometric structure by modeling anchor distributions, thereby enabling adaptive similarity measurement. Experimental results demonstrate that the proposed method efficiently and accurately characterizes context-dependent semantic similarity, yielding substantial improvements in retrieval performance.
In zero-shot learning, manually defined class-level semantic prototypes suffer from instance-level misalignment (e.g., occlusion, viewpoint variation) and coarse-grained semantic imprecision, leading to biased vision–semantic mapping and degraded knowledge transfer to unseen classes. To address this, we propose a prototype-guided curriculum learning framework: (1) samples are progressively selected for training based on cosine similarity to assess visual–semantic alignment; (2) an instance-level feedback mechanism dynamically refines class-level prototypes, mitigating the impact of label noise. Our method requires no additional supervision and significantly enhances model generalization. Extensive experiments on AWA2, SUN, and CUB benchmarks demonstrate state-of-the-art classification accuracy, confirming both effectiveness and robustness.
Existing prototype-based self-supervised learning relies on a single prototype to represent all features within a cluster, failing to capture semantic diversity in data space. This work proposes Self-Organizing Prototypes (SOP), which abandons fixed prototypes and instead dynamically organizes multiple semantically similar support embeddings (SEs) to collaboratively model local feature structures. Methodologically: (i) it introduces the first multi-prototype collaborative representation mechanism; (ii) it designs a non-parametric SOP-MIM masked modeling task; and (iii) it integrates non-parametric contrastive learning, reconstruction loss, and dynamic SE organization for fully parameter-free feature-space modeling. SOP achieves state-of-the-art performance across diverse downstream tasks—including image retrieval, linear evaluation, fine-tuning, and object detection—with particularly pronounced gains when adapted to large-scale models.
This work addresses the challenges of category confusion due to inter-class similarity and the lack of spatial localization in similarity matching within few-shot object detection. To this end, the authors propose Text-anchored Semantic Masks (TSMa) and a Stage-aligned Hierarchical Autoregressive Regression (SHARe) mechanism. TSMa leverages textual features as semantic anchors to suppress style-induced interference and enhance intrinsic class discriminability. SHARe formulates bounding box regression as a multi-stage progressive refinement process, aligning the abstraction capabilities of different Vision Transformer layers with corresponding regression stages to improve localization accuracy. Notably, the method generalizes to novel categories without additional training and achieves a new state-of-the-art performance on the COCO few-shot detection benchmark, surpassing the previous best approach by 10.1 nAP.
Existing prototype alignment methods in heterogeneous federated learning enforce clients with diverse architectures to align within a unified feature subspace, thereby constraining model expressiveness. This work proposes FedSAF, a novel structural alignment paradigm that shifts the alignment objective from coordinate-wise matching to preserving the consistency of inter-class relational structures. By decoupling semantic structure alignment from shared feature bases, FedSAF models class relationships through prototypes and integrates them into a distributed optimization framework. Extensive experiments demonstrate that FedSAF significantly outperforms current heterogeneous federated learning approaches across multiple benchmarks, achieving accuracy improvements of up to 3.52%.
Concept alignment lacks a unified definition, and existing methods optimize divergent objectives under the same terminology, obscuring its fundamental nature. This work formalizes its multidimensional structure by decomposing it along two axes—“alignment targets” and “alignment levels”—and identifies four distinct alignment properties, revealing that current approaches satisfy only subsets of these. To address this limitation, we propose Coupled Sparse Autoencoders (CoSAE), a framework that jointly optimizes multiple alignment objectives, alongside InterVenchA, an interventional evaluation benchmark. Experiments demonstrate that optimizing a single objective fails to reliably recover other alignment properties, whereas CoSAE achieves strong instance-level conceptual consistency using merely 0.1% paired data.
This work addresses the challenge of constructing modular AI systems from independently trained neural models, whose latent spaces are often incompatible. To overcome this, the authors propose an enhanced relative representation framework that leverages learnable semantic prototypes as cross-model anchors and replaces cosine similarity with a whitened inner product that preserves magnitude information while being invariant to affine transformations, thereby enabling geometry-aware similarity measurement. This approach effectively aligns embedding spaces across heterogeneous architectures—including small language models—and achieves near-lossless information transfer and stable zero-shot communication in vision-and-language tasks, significantly improving cross-model representational consistency.
Existing layout-to-image generation methods suffer from fragmented representations under few-shot, atypical scenarios, leading to image distortion and loss of detail. This work proposes a representation-driven framework that explicitly decouples semantic identity from visual primitives for the first time. Semantic anchoring aggregates category-level semantics to stabilize object identity, while primitive injection models recombinable local primitives to enhance fine-grained details. Furthermore, a concept-guided mechanism incorporating saliency-aware optimization is introduced to improve foreground semantic consistency. Evaluated under a strict 5-shot setting, the proposed method consistently outperforms state-of-the-art approaches across multiple atypical domains, achieving notable improvements in both visual fidelity and layout alignment.