BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
High-precision 3D medical image segmentation typically relies on dense multi-scale feature maps, resulting in substantial memory consumption and computational overhead. To address this, this work proposes the BATS architecture, which employs a boundary-aware hybrid-resolution token mechanism to preserve fine-grained processing near class boundaries while using coarse representations elsewhere, thereby constructing an input-adaptive sparse hierarchical structure that is subsequently reconstructed into a dense segmentation output. The method introduces a novel boundary-correlation-independent prediction mechanism to prevent coarse-level misclassifications from suppressing fine details and incorporates a parent-cluster attention module for efficient cross-scale contextual fusion—eliminating the need for dense feature maps or neighborhood search. Evaluated on five CT/MRI datasets, BATS achieves segmentation performance only 0.37 Dice points below the strongest baseline while reducing peak GPU memory usage by over 53% and accelerating inference by up to 30% on KiTS and LiTS benchmarks.
📝 Abstract
Many high-performing volumetric segmentation models maintain dense multi-scale feature maps, leading to high activation memory and inference cost. We present BATS (Boundary-Aware Token Selection), a 3D medical image segmentation architecture that concentrates fine-resolution processing near predicted class boundaries. A dense boundary predictor identifies where additional resolution is needed, while a fine-first context cascade constructs an input-dependent mixed-resolution hierarchy. Homogeneous regions are represented coarsely, with finer tokens retained around boundaries, thin structures, and small targets. The sparse hierarchy is refined and rasterised into a dense segmentation. BATS predicts boundary relevance independently at every resolution level, preventing an erroneous coarse-scale decision from suppressing fine-scale evidence. Parent cluster attention further injects hierarchical ancestor tokens into local attention neighbourhoods, providing cross-scale context without dense multi-scale feature maps or cross-scale neighbour search. We evaluate BATS on five public CT and MRI datasets using the standardised nnU-Net Revisited protocol. BATS achieves the highest LiTS Dice among the compared methods and averages within 0.37 Dice points of the strongest dense baseline, MedNeXt-L, across the five datasets. Relative to MedNeXt-L, it reduces peak allocated GPU memory by more than 53% on KiTS, LiTS, and BraTS. Inference is up to 30% faster on KiTS and LiTS, which retain fewer tokens, but slower on the more token-dense BraTS. Mixed-resolution processing therefore provides consistent memory savings, while runtime and accuracy gains depend on dataset boundary density.
Problem

Research questions and friction points this paper is trying to address.

volumetric segmentation
activation memory
inference cost
3D medical image
multi-scale feature maps
Innovation

Methods, ideas, or system contributions that make the work stand out.

boundary-aware token selection
mixed-resolution processing
sparse hierarchical segmentation
parent cluster attention
resource-efficient 3D segmentation
🔎 Similar Papers
No similar papers found.