Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics
This study investigates whether the internal mechanisms of the scientific foundation model Walrus align with physical principles when reproducing continuum dynamics, and examines the relationship between its representations and performance. By introducing sparse autoencoders (SAEs) at specific layers, the work pioneers the use of enstrophy—the integral of squared vorticity—for physically grounded filtering and prioritization of large-scale features, complemented by comparative numerical simulations. The findings reveal that while the model’s feature activations exhibit segment-wise consistency, they do not correspond to physically meaningful decompositions. Notably, certain output inaccuracies, such as excessive energy dissipation, can be traced to variations in specific SAE features. The study underscores fundamental challenges in achieving representational fidelity and interpretability in scientific foundation models.