🤖 AI Summary
This study addresses the resolution versus long-horizon trade-off caused by quadratic attention costs in Vision Transformers for solving partial differential equations, alongside the lack of dynamic adaptation mechanisms. We propose WAMRViT, the first machine learning surrogate natively supporting multi-level adaptive mesh refinement (AMR) data. By integrating dynamic quadtree tokenization, wavelet-inspired refinement criteria, and 3D rotary position embeddings, it overcomes fixed token-budget constraints to achieve fully adaptive topological processing without predefined token counts. Experiments demonstrate that WAMRViT significantly improves region-of-interest accuracy and long-horizon rollout stability on both uniform and multiscale grids using fewer tokens. Furthermore, it effectively balances fine-grained error suppression with computational efficiency in complex combustion problems.
📝 Abstract
The quadratic attention cost of Vision Transformers (ViTs) forces a trade-off between spatial resolution and rollout horizon, particularly for fine-scale PDEs where shocks, reaction fronts, and material interfaces occupy small, evolving regions of the domain. Conventional neural surrogates also lack mechanisms to adapt resolution dynamically. We propose WAMRViT, a ViT that tokenizes inputs as balanced quadtrees using a wavelet-inspired refinement criterion, jointly encodes position and refinement level with 3D rotary positional embeddings, and regrids in cell space during inference for stable long-horizon rollouts. A multi-scale variant retains each leaf at its native source resolution and lets the model learn across resolution levels. Unlike fixed-budget adaptive-tokenization methods, WAMRViT imposes no predetermined token count and supports fully adaptive topology throughout autoregressive rollout. To our knowledge, it is the first machine-learning surrogate to natively tokenize multi-level Adaptive Mesh Refinement (AMR) data. On uniform-grid benchmarks, uniform-patch WAMRViT improves finest-level region-of-interest VRMSE over a finest-patch uniform ViT while using substantially fewer tokens. The multi-scale variant achieves the lowest first-step full-field RMSE and VRMSE on both benchmarks and improves rollout-averaged full-field and refined-region accuracy at long horizons. With parallelized regridding, its end-to-end rollout cost lies between finest-patch and approximately token-matched coarser-patch ViTs. On a complex AMR combustion problem whose finest features cannot be represented natively by the evaluated uniform-grid baselines, WAMRViT operates directly on adaptive cells and substantially reduces finest-level error at matched transformer capacity. Code: https://github.com/tonyzyl/wamrvit