Institution profile

British Antarctic Survey

Academic institutioneurope · gb
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Warping Earth Observations for better ice labeling in the Marginal Marginal Ice Zone

Aug 12, 2026

This study addresses the challenge of pixel-level misalignment between multi-source satellite images in dynamic marginal ice zones, which hinders effective multimodal fusion and accurate sea ice segmentation. To overcome this, the authors propose a novel mutual information-based warping framework—the first to apply this approach for spatial alignment between Sentinel-1 synthetic aperture radar and MODIS visible/thermal infrared imagery. Integrated with sparse point annotations, the method enables high-precision dense semantic segmentation, departing from conventional training paradigms that rely on coarse-grained ice charts. Evaluated on a dataset comprising 43 scenes with 2,088 pixel-level labels derived from 7,046 expert-annotated points, experiments demonstrate significantly improved segmentation accuracy after alignment, effectively achieving robust generalization from sparse annotations to dense predictions.

0 citationsRead paper

ConSens: Assessing context grounding in open-book question answering

Apr 30, 2025

In open-book question answering, existing evaluation methods suffer from bias, poor scalability, and reliance on external systems, hindering accurate measurement of a model’s contextual dependency. To address this, we propose ConSens—a novel metric that quantifies a language model’s “context anchoring ability” by measuring the relative perplexity difference between context-aware and context-agnostic generations—using the model itself as both evaluator and generator. ConSens requires no fine-tuning, eliminates dependence on external judge models, and supports zero-shot, cross-model evaluation with high interpretability. Extensive experiments across multiple datasets demonstrate that ConSens achieves strong correlation with human judgments (Spearman’s ρ > 0.85), significantly outperforms baselines such as LLM-as-a-judge in discriminative power, and reduces computational overhead by over 90%.

0 citationsRead paper
Recent publications

Latest Papers

Warping Earth Observations for better ice labeling in the Marginal Marginal Ice Zone

Aug 12, 2026

This study addresses the challenge of pixel-level misalignment between multi-source satellite images in dynamic marginal ice zones, which hinders effective multimodal fusion and accurate sea ice segmentation. To overcome this, the authors propose a novel mutual information-based warping framework—the first to apply this approach for spatial alignment between Sentinel-1 synthetic aperture radar and MODIS visible/thermal infrared imagery. Integrated with sparse point annotations, the method enables high-precision dense semantic segmentation, departing from conventional training paradigms that rely on coarse-grained ice charts. Evaluated on a dataset comprising 43 scenes with 2,088 pixel-level labels derived from 7,046 expert-annotated points, experiments demonstrate significantly improved segmentation accuracy after alignment, effectively achieving robust generalization from sparse annotations to dense predictions.

0 citationsRead paper

ConSens: Assessing context grounding in open-book question answering

Apr 30, 2025

In open-book question answering, existing evaluation methods suffer from bias, poor scalability, and reliance on external systems, hindering accurate measurement of a model’s contextual dependency. To address this, we propose ConSens—a novel metric that quantifies a language model’s “context anchoring ability” by measuring the relative perplexity difference between context-aware and context-agnostic generations—using the model itself as both evaluator and generator. ConSens requires no fine-tuning, eliminates dependence on external judge models, and supports zero-shot, cross-model evaluation with high interpretability. Extensive experiments across multiple datasets demonstrate that ConSens achieves strong correlation with human judgments (Spearman’s ρ > 0.85), significantly outperforms baselines such as LLM-as-a-judge in discriminative power, and reduces computational overhead by over 90%.

0 citationsRead paper