🤖 AI Summary
This study addresses the difficulty of disentangling structural contributions from model capacity in mask refinement, as well as its poor cross-generator generalization. We formulate the refinement process as conditional random field (CRF) inference. Specifically, a zero-state mechanism is introduced to handle weakly labeled regions, decoupling boundary from non-boundary errors. Furthermore, DINOv2 features are leveraged to construct image-conditioned latent region consistency constraints, and damped mean-field iterations are unrolled to perform structured corrections. This approach establishes a verification paradigm demonstrating that explicit structure outperforms black-box capacity. Extensive experiments on datasets such as COD10K show that our method significantly surpasses parameter-matched baselines, achieves performance comparable to foundation models with minimal parameters, and effectively enhances generalization to unseen generators.
📝 Abstract
Learned mask refiners improve segmentation accuracy, but it is hard to tell how much of the improvement comes from explicit structure rather than from extra capacity, and whether it holds up when the mask generator or its error distribution changes. UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables. Its energy has a corrected unary term, learned local pairwise interactions, and image-conditioned latent-region consistency, with a null state that lets a region with weak label agreement withdraw from the consistency term; inference unrolls damped mean-field updates on this one energy. To isolate the effect of structure, we compare against recurrent black-box refiners that read the same inputs and receive the same parameter budget, stage count, and supervision. On COD10K, UnfoldCRF beats the strongest matched control by 1.0 $F^\omega_\beta$ point, improves all four COD metrics, and lowers the fraction of images made worse from 11.7\% to 8.5\%. Under a train-once protocol over five datasets and several mask sources, the 2.6M-parameter variant gains 4.2 mean $\Delta$IoU against 2.0 for its control, and a variant built on frozen DINOv2 features matches the strongest foundation-model refiner with about a seventh of its resident parameters while staying ahead of its own control. On mask generators never seen in training, the gain is 2.0 $F^\omega_\beta$ points against 0.6 for the control. Zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors. Code and supporting materials will be publicly released.