🤖 AI Summary
This work addresses the inconsistency in existing data augmentation methods when jointly transforming images and their associated multimodal annotations—such as masks, bounding boxes, and keypoints—where mismatched random transformations often lead to misaligned training samples and degraded data quality. To resolve this, the authors propose a unified augmentation framework that encapsulates the augmentation pipeline into composable Compose objects, rigorously synchronizing transformation parameters and random seeds across all modalities. The framework supports diverse data types including images, masks, bounding boxes, keypoints, stereo views, video frames, and volumetric data. Furthermore, it incorporates an augmentation history logging and replay mechanism, ensuring fully reproducible and traceable augmentation processes. This approach significantly enhances the reliability of training data and improves model robustness.
📝 Abstract
Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume. Code paths that choose these values separately can silently misalign the data.
AlbumentationsX keeps the transform list, probabilities, annotation settings, and random seed in one Compose object. Each call chooses random values once and applies them to every supported part of the training example. The library keeps each object's mask, box, and label together and lets projects add their own transforms. It can also save the pipeline definition, show what happened in one call, and run that call again.
The examples place Compose after files have been decoded into arrays and before PyTorch groups examples into a batch. AlbumentationsX executes the declared transforms. Practitioners still decide whether a flip, crop, color change, or other operation preserves the correct label for their task.