Score
Designs and implements test-time augmentation pipelines that rescale and flip input images (and apply stochastic augmentations) to run a fixed-size model multiple times and aggregate the per-augmentation predictions (ensembling) into a single, more robust output. Also develops methods to detect unstable or augmentation-sensitive predictions and performs analyses such as timescale-separation to diagnose and restore the benefits of multi-scale inference.
This study systematically evaluates the effectiveness of test-time augmentation (TTA) in medical image classification and finds that, contrary to common assumptions, TTA often significantly degrades accuracy—by as much as 31.6 percentage points—across most scenarios, with only marginal gains observed in specific dermatological tasks (+1.6%). Leveraging the MedMNIST v2 benchmark, four model scales, and multiple standard TTA strategies, the authors identify the primary cause of performance degradation as a mismatch between training and test distribution, exacerbated by incompatible batch normalization statistics. Notably, the work demonstrates for the first time that intensity-based transformations consistently outperform geometric ones. These findings challenge the prevailing assumption of TTA’s universal efficacy and underscore the necessity of carefully selecting augmentation strategies in medical imaging applications.
Conformal prediction often yields overly large and uninformative prediction sets, struggling to balance statistical validity with practical utility. To address this, we propose the first framework that systematically integrates test-time augmentation (TTA) into conformal prediction, introducing lightweight inductive biases during inference—without requiring model retraining—and maintaining compatibility with arbitrary conformal score functions (e.g., APS, RAPS) and adaptive quantile calibration. Our method preserves rigorous marginal coverage guarantees while substantially improving set informativeness and compactness. Extensive evaluation across three diverse datasets, three model architectures, and multiple distribution shifts demonstrates consistent improvements: average prediction set size is reduced by 10–14%, validating the framework’s generality and robustness.
This study addresses the high computational costs of data augmentation ensembles and their inefficiency in leveraging task symmetries by proposing Stochastic Weight Averaging (SWA) as a replacement for repetitive ensembling. Through approximation analysis via the Ornstein-Uhlenbeck process, we reveal that SWA enhances model equivariance beyond conventional performance gains in the infinite-width limit. Experiments on visual and graph classification tasks demonstrate the method’s superiority across both discrete and continuous symmetries. These findings validate SWA as an effective alternative to traditional ensembling, providing new theoretical foundations and a practical paradigm for efficiently exploiting data augmentation. This work thus bridges the gap between computational efficiency and symmetry-aware learning, offering significant implications for scalable representation learning in structured domains.
This work addresses the poor robustness and unstable online updates of test-time adaptation (TTA) under adversarially corrupted test streams by proposing SAFER, a training-free robust TTA framework. SAFER constructs stable predictions through reliability-guided stochastic augmentation and correlation-weighted pooling, and introduces an adaptive ensembling strategy based on feature divergence to simultaneously enhance adversarial robustness and preserve performance on clean samples. As the first systematic study of robust TTA in the adversarial streaming setting, SAFER significantly improves the robustness of diverse TTA methods against PGD attacks on PACS, VLCS, and OfficeHome datasets while maintaining competitive accuracy on clean data.
This work addresses the challenge of improving the accuracy of large language models under fixed inference compute budgets, rather than merely increasing output diversity. It proposes a test-time input augmentation (TTA) strategy that systematically explores three classes of input-side enhancements—semantic paraphrasing, lexical perturbation, and visual transformation—combined with chain-of-thought prompting and prediction aggregation. Empirical results demonstrate, for the first time, that this approach significantly outperforms conventional self-consistency on five out of six benchmark tasks, achieving statistically significant accuracy gains at comparable computational cost. The method yields an average cost-effectiveness improvement of approximately 1.8×, exhibits particular efficacy for medium-scale models, and achieves Pareto dominance in the trade-off between cost and performance.
This study addresses the limited out-of-distribution (OOD) generalization of tabular data after image-based encoding under distribution shift. It presents the first systematic evaluation of test-time augmentation (TTA) across multiple tabular-to-image transformation methods, including TINTO and DeepInsight. Leveraging the TableShift benchmark, the authors combine six encoding strategies with 25 TTA techniques spanning geometric, photometric, structural, frequency-domain, Mixup, and composite categories. Their analysis reveals that composite and photometric augmentations consistently enhance OOD performance, whereas frequency-domain transformations generally degrade it. This work demonstrates the robustness potential of TTA in tabular image classification and provides practical, low-variance augmentation protocols for real-world deployment.
本文通过文献综述和大规模实证研究,评估了50种图像增强技术作为深度学习图像检索系统的测试生成方法,以提高系统可靠性。
This work addresses the inconsistency in existing data augmentation methods when jointly transforming images and their associated multimodal annotations—such as masks, bounding boxes, and keypoints—where mismatched random transformations often lead to misaligned training samples and degraded data quality. To resolve this, the authors propose a unified augmentation framework that encapsulates the augmentation pipeline into composable Compose objects, rigorously synchronizing transformation parameters and random seeds across all modalities. The framework supports diverse data types including images, masks, bounding boxes, keypoints, stereo views, video frames, and volumetric data. Furthermore, it incorporates an augmentation history logging and replay mechanism, ensuring fully reproducible and traceable augmentation processes. This approach significantly enhances the reliability of training data and improves model robustness.
研究比较了五个库的七种输入路径,从RGB JPEG文件到CUDA float16批处理,评估图像增强管道的吞吐量和GPU内存使用,以优化数据预处理效率。
研究通过LIBERO-CTRL六轴基准测试,探讨了单轴评估能否推断复合鲁棒性,并揭示了仅凭单轴成功率无法完全表征复合鲁棒性的现象。