🤖 AI Summary
This study investigates whether the Transformer attention mechanism is essential in Transolver and identifies its performance bottlenecks. Through systematic ablation experiments, we reveal that global token mixing, rather than attention itself, constitutes the core mechanism. Grounded in averaging theory, we prove that slice and anti-slice operations combined with a multilayer perceptron (MLP) suffice to achieve universal approximation, and accordingly propose FlashSlice, an efficient computational kernel. This work provides the first demonstration that Transolver operates effectively without a Transformer architecture. By eliminating attention while preserving state-of-the-art accuracy, our approach substantially reduces memory consumption and computational overhead, offering a minimalist yet highly efficient alternative for neural operator design.
📝 Abstract
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empirical and theoretical analysis to elucidate the mechanisms which are responsible for model performance. To this end, we perform careful ablations on a challenging suite of nine 3D fluid dynamics benchmarks to find that replacing token attention with a constant linear map does not affect the accuracy. Thus, Transolver does not need a Transformer at all. However, removing the global mixing (slicing/deslicing) or doing it only once leads to performance collapse. We leverage the theory of averaging neural operators to explain and corroborate our findings by showing that just slicing/deslicing, in conjunction with pointwise MLPs, already suffices for universal approximation of continuous operators and attention is redundant in this context. Finally, we provide a novel FlashAttention-style efficient implementation of the key slicing/deslicing module of Transolver. This flashslice kernel streams slice and deslice over the points without materializing the heavy slice-weight tensor, while reproducing the best available implementation to floating point error. At the same time, it leads to very significant memory and compute savings, particularly at large slice counts.