Score
Designs and implements contrastive pretraining objectives and augmentation pipelines that explicitly model, preserve, or align phase information in signals so positive and negative pairs reflect phase-aware similarity. Trains encoders with those phase-aware contrastive losses to produce discriminative, phase-sensitive representations of signals (e.g., audio or other phase-containing time-series) that improve downstream classification or detection models.
This study investigates the human perceptual threshold of spectrum-agnostic Phase Offset Distortion (POD), rigorously establishing—via controlled psychoacoustic experiments—that POD is highly imperceptible. Building on this finding, we propose a theoretically grounded, lossless, and semantics-preserving POD-based data augmentation paradigm: POD is modeled as a learnable perturbation in the phase domain and integrated into end-to-end training of audio models (CNNs and Transformers). Evaluated on speech recognition and environmental sound classification, our method yields an average accuracy improvement of 2.3%, significantly mitigates overfitting, and strictly preserves original audio fidelity and semantic content. Our core contributions are twofold: (1) the first empirical validation of POD’s perceptual robustness, and (2) the introduction of the first phase-domain data augmentation framework with formal theoretical guarantees of losslessness and semantic invariance.
Neural networks exhibit high sensitivity to semantically irrelevant transformations—such as ECG phase shifts or IMU sensor rotations—leading to representation degradation and performance collapse. To address this, we propose a structured contrastive learning framework that pioneers the disentanglement of latent representations into three functionally distinct subspaces: invariant (encoding semantics), variant (modeling controlled transformations), and free (capturing residual variability). This design jointly ensures semantic invariance and explicit modeling of admissible variations, unifying robustness with interpretability. Our method requires no architectural modifications; instead, it achieves end-to-end structural learning via semantic grouping and embedded training, advancing contrastive learning from passive data augmentation toward active structural modeling. Evaluated on ECG phase-invariance tasks, our approach achieves a similarity score of 0.91 (+0.66 improvement); on IMU pose-robust activity recognition, it attains 86.65% accuracy and 95.38% rotation consistency.
Although contrastive decoding (CD) has demonstrated performance improvements in large audio language models (LALMs), its underlying mechanisms and conditions for effectiveness remain unclear. This work systematically evaluates four CD strategies across diverse LALM architectures and introduces a transition-matrix-based framework to analyze error patterns. The analysis reveals that CD effectively corrects only two error types—“false negatives” (misclassifying audio as absent) and “uncertain guesses”—but fails to address erroneous reasoning or confidently incorrect predictions. Among the evaluated approaches, Audio-Aware Decoding and Audio Contrastive Decoding emerge as the most effective strategies. Building on these findings, the study establishes principled guidelines for selecting CD methods based on the baseline model’s characteristic error profiles, offering both theoretical insights and practical recommendations for optimizing LALM inference.
Class imbalance significantly degrades the performance of contrastive learning, yet the underlying mechanisms by which it affects training dynamics and induces representation bias remain theoretically underexplored. This work addresses this gap by introducing a novel theoretical framework that analyzes the training process of Transformer-based contrastive learning under imbalanced data through the lens of neuron weight evolution. The analysis reveals three characteristic phases in the evolution of neuronal weights during training. Guided by these theoretical insights, the authors propose a targeted neuron pruning strategy that effectively mitigates representation bias. Experimental results demonstrate that the proposed method substantially enhances feature separability and overall representation quality in imbalanced scenarios.
Time-series behavioral biological data pose challenges for contrastive learning due to reliance on manual hyperparameter tuning and high computational cost in hand-crafted view generation. Method: This paper proposes LEAVES, a learnable automatic augmentation strategy module that introduces the first differentiable view generation mechanism into time-series contrastive learning frameworks. Leveraging adversarial training, LEAVES jointly optimizes augmentation hyperparameters in an end-to-end manner to adaptively learn optimal augmentation policies without human intervention. Contribution/Results: LEAVES significantly improves view plausibility and representation discriminability. Evaluated on multiple standard time-series benchmarks, it outperforms both manually tuned baselines and state-of-the-art contrastive learning methods (e.g., TS-TCC variants), achieving average accuracy gains of 3.2–5.8% across downstream tasks including classification and forecasting.
This work addresses a critical limitation in self-supervised dynamic representation learning, where existing contrastive predictive objectives often misinterpret slowly varying noise within trajectories as genuine dynamical signals, leading to noise-dominated representations and degraded downstream performance. The authors identify this issue as stemming from an inherent inductive bias flaw in standard contrastive objectives and propose a general corrective principle: sampling negative examples from within the same trajectory to eliminate predictive shortcuts introduced by slow-varying noise, thereby compelling the encoder to focus on the true dynamical variables governing system evolution. Experiments based on frameworks such as JEPA and DySIB on synthetic moving-point and rigid-pendulum video datasets demonstrate that the proposed approach effectively disentangles slow noise from authentic dynamics, yields representations whose quality improves with trajectory length, and significantly enhances downstream task performance under strong noise conditions.
This work investigates why contrastive learning yields effective representations under simple image augmentations. By analytically characterizing the optimal solution of the contrastive loss, the study theoretically establishes—for the first time—that, under specific augmentations, the optimal first-layer filters are sinusoidal functions. This insight leads to a derived CNN architecture comprising sinusoidal filters, pointwise nonlinearities, global average pooling, and a partially whitening linear layer. The authors further introduce a water-filling algorithm based on the data’s power spectrum to compute the frequencies and weights of these filters. Experiments across multiple image datasets and augmentation strategies confirm that the first layer indeed learns sinusoidal filters and performs partial whitening, thereby revealing the underlying mechanism by which contrastive learning shapes its representations.
本文提出CoJEPA方法,结合对比学习和JEPA解决音乐表示中的全局-局部问题,通过共享骨干网络联合训练以获得更丰富的音乐表示。
This work addresses the limitations of handcrafted augmentation strategies in time series contrastive learning, which often introduce spurious correlations and suffer from poor generalization. To overcome these issues, the authors propose a novel paradigm that explicitly encodes temporal shift invariance to construct deterministic views, thereby replacing conventional domain-knowledge-dependent augmentations. This approach leverages temporal shift invariance alone to generate effective positive and negative sample pairs, significantly reducing reliance on manual intervention. Evaluated across six real-world benchmarks and the UCR/UEA archive, the method achieves state-of-the-art performance while substantially accelerating training. Furthermore, the study systematically investigates the impact of batch size and the number of negative samples on model effectiveness, offering valuable insights into the design of contrastive learning frameworks for time series data.
This work addresses the susceptibility of large audio language models to hallucinations caused by linguistic priors overpowering acoustic evidence. To mitigate this, the authors propose a task- and sample-adaptive perturbation selection mechanism within a contrastive decoding framework. Leveraging a structured audio perturbation bank spanning temporal, spectral, frequency, and amplitude domains, the method dynamically selects optimal negative-sample perturbation strategies and employs a lightweight selector for efficient routing. The approach yields a 4.3% absolute improvement in accuracy on existence tasks and significantly boosts performance on temporal tasks from 74.7% to 81.4%. Furthermore, the study validates the efficacy of binary-constrained prompts, underscoring the critical role of adaptive perturbation strategies in alleviating hallucinations in audio language models.