Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping

๐Ÿ“… 2026-08-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the need for explainability and environmental robustness in AI-assisted navigation for Arctic shipping by proposing a reward modeling approach based on inverse reinforcement learning (IRL). The authors systematically compare linear shared reward models, nonlinear shared reward models, and a meta-IRL model incorporating vessel-class latent variables, using preregistered feature-masking ablation experiments and multidimensional evaluation metricsโ€”including predictive accuracy, route fidelity, and reward transferability. Results show that the nonlinear reward model improves held-out likelihood by 50.9% over the linear baseline, while introducing latent variables degrades performance by 16.5%. This indicates that heterogeneity in vessel behavior is primarily driven by observable route and environmental factors rather than unobserved preferences.
๐Ÿ“ Abstract
Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments. Inverse reinforcement learning (IRL) provides a framework for recovering such rewards from vessel trajectories, while recent meta-IRL methods introduce latent context variables to capture behavioral heterogeneity. However, it remains unclear whether these latent representations recover genuinely hidden preferences or simply re-encode information already available in the observed state. We conduct a controlled evaluation on 3,186 AIS-derived voyages from 202 vessels across nine Arctic shipping seasons, comparing a linear shared reward, a nonlinear shared reward, and a latent-context model built on the same nonlinear architecture. The nonlinear reward improves held-out likelihood by 50.9% over the linear baseline, whereas adding vessel-specific latent context reduces performance by 16.5%. Behavioral analysis, context probes, and a pre-registered feature-hiding ablation show that apparent vessel-level variation is largely explained by observable route and environmental conditions rather than hidden vessel-specific factors. Moreover, predictive accuracy, route fidelity, and reward transfer yield different model rankings, demonstrating that no single metric is sufficient to evaluate learned rewards. These findings motivate testing whether the observed route, environmental, and vessel features already explain behavioral variation before adding per-vessel latent context. This supports more trustworthy AI deployment in safety-critical domains.
Problem

Research questions and friction points this paper is trying to address.

inverse reinforcement learning
latent context
Arctic shipping
behavioral heterogeneity
reward modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

inverse reinforcement learning
latent context
Arctic shipping
behavioral heterogeneity
controlled evaluation
V
Vaishnav Vaidheeswaran
Faculty of Computer Science, Dalhousie University, Halifax, NS, Canada
D
Dilith Jayakody
Faculty of Computer Science, Dalhousie University, Halifax, NS, Canada
B
Biruk Ambaw
Faculty of Computer Science, Dalhousie University, Halifax, NS, Canada
J
Jaswanth Kumar
Department of Industrial Engineering, Faculty of Engineering, Dalhousie University, Halifax, NS, Canada
M
Md Mahbub Alam
Faculty of Computer Science, Dalhousie University, Halifax, NS, Canada
Gabriel Spadon
Gabriel Spadon
Assistant Professor, Faculty of Computer Science, Dalhousie University
Data MiningMachine LearningNetwork ScienceGeoinformatics