Mechanistic Interpretability of Atmospheric Rivers in GraphCast

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited interpretability of AI-based weather models, whose internal representations of atmospheric information and computational logic remain largely opaque to conventional analysis. To investigate this, we apply standard and Matryoshka Sparse Autoencoders (SAEs) to dissect the GraphCast model, specifically examining its internal mechanistic representations of atmospheric rivers. Our findings reveal that the model stably encodes integrated vapor transport (IVT)β€”a variable neither provided as input nor targeted during trainingβ€”as a persistent internal feature. Furthermore, the Matryoshka SAE effectively ranks these learned concepts and elucidates their interrelationships. Causal intervention analyses confirm that atmospheric river representations persist across network depth and exert genuine causal influence on model predictions. This work establishes a novel methodological pathway for enhancing the interpretability of AI-driven meteorological models, which is increasingly critical in the context of climate change.
πŸ“ Abstract
While AI weather models now rival operational forecasts, how they represent the atmosphere internally remains an open question: feature attribution reveals which input patterns matter, not what the model computes or how it combines information internally. We train sparse autoencoders (SAEs) on GraphCast to uncover its learned concepts, using atmospheric rivers as our phenomenon of focus. Both standard and Matryoshka SAEs show GraphCast computes atmospheric river intensity, measured by integrated vapor transport (IVT), as a stable internal variable, despite IVT being neither an input nor a target. In contrast to the unstructured concept retrieval of the standard SAE, the Matryoshka SAE orders concepts by importance and exposes their relations. Atmospheric river concepts persist across depth and direct interventions confirm causality. This method offers a way to find internal variables and determine which of them the model actually relies on, which is a prerequisite for asking whether those variables remain meaningful as the phenomenon changes under a warming climate.
Problem

Research questions and friction points this paper is trying to address.

Mechanistic Interpretability
GraphCast
Atmospheric Rivers
Internal Representations
AI Weather Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mechanistic Interpretability
Sparse Autoencoders
GraphCast
Atmospheric Rivers
Matryoshka SAE
πŸ”Ž Similar Papers
No similar papers found.