🤖 AI Summary
This study investigates whether text-conditioned neural PDE surrogates genuinely exploit semantic information or merely reduce numerical errors. Building upon the Fourier Neural Operator architecture with Feature-wise Linear Modulation, this work proposes a path-control diagnostic framework and an OperatorCLIP evaluation metric. By systematically comparing unconditional, constant-sentence, and task-text models across diverse fluid dynamics tasks, it examines the efficacy of text guidance. The findings reveal a methodological pitfall wherein InfoNCE loss fails to identify matching pairs under single-description training. Furthermore, constant conditioning yields lower errors without exhibiting reliable semantic ranking, indicating that current paradigms do not effectively leverage textual semantics. These results underscore the necessity of rigorous controlled experiments to validate the genuine utility of text in neural PDE modeling.
📝 Abstract
Lower error from a text-conditioned neural surrogate does not, by itself, show that the model uses the meaning of the text. We examine this attribution problem with OperatorCLIP, comparing an unconditioned FNO, a constant-sentence FiLM control, and a fixed task description trained with contrastive alignment. Three-seed experiments cover Darcy2D, ShallowWater2D, and three-dimensional compressible Navier-Stokes (CNS3D). Constant conditioning has lower mean test error on both 2D tasks. Relative to this control, task text plus alignment has a similar mean on ShallowWater2D and CNS3D and a higher mean on Darcy2D; these descriptive comparisons have substantial seed uncertainty. The latter comparison changes both prompt content and loss, so it isolates neither effect. The text encoder is trained from scratch, and each conditioned model sees only one description during training. In this regime, pairwise InfoNCE cannot identify matched pairs and has minimum $\log B$. Prompt interventions show no reliable semantic ordering. This methodological caution demonstrates why pathway controls are needed; it neither establishes semantic competence of the encoder nor tests the effectiveness of text under varying physical context.