Atomizer-IO: Beyond Pixels, Patches and Grids
This study addresses the limitation of existing vision architectures that rely on regular grid assumptions, rendering them ineffective for unstructured sensor data. To overcome this, we propose a novel architecture based on atomic representations, introducing an "observation-first" paradigm. By leveraging measurement metadata and anchor-based local cross-attention mechanisms, our approach infers structural information directly from physical relationships, thereby eliminating grid constraints entirely. The proposed architecture seamlessly generalizes to diverse inputs, such as unordered 3D point clouds, without requiring task-specific redesigns. Furthermore, it achieves performance levels comparable to specialized models across multiple tasks while enabling post-training control over inference costs. These results comprehensively validate the versatility and flexibility of the proposed grid-free interface for processing heterogeneous sensory data.