iterative region refinement

Designs and builds algorithms and pipelines that iteratively locate, score (triage), and refine regions of interest within an input instance, using region-wise importance or triage scores to focus computation and progressively improve region-level representations, boundaries, or feature maps. Implements repeated region-focused inspection and progressive refinement passes so flagged areas receive higher representational capacity and successive updates until a stopping criterion is met.

iterativeregionrefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.8
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of detecting small and diverse tumor lesions in large-scale CT imaging, where extreme foreground-background imbalance and high false-positive rates hinder performance. To this end, we propose GF-Screen, a novel framework that, for the first time, effectively applies reinforcement learning to pan-cancer screening by emulating radiologists’ “glance-and-focus” strategy: a Glance model coarsely identifies suspicious subregions, while a Focus model performs precise lesion segmentation, with the segmentation outcome serving as a reward signal to guide the Glance model’s region selection. We further introduce a group-wise relative learning paradigm that prioritizes high-advantage predictions within each group to enhance efficiency and suppress false positives. Evaluated on 16 internal and 7 external datasets covering nine lesion types, our method achieves top rank on the MICCAI FLARE25 Pan-Cancer Challenge leaderboard, surpassing the FLARE24 champion by 25.6% in DSC and 28.2% in NSD.

false positivesforeground-background imbalancelarge-scale CT scans

Vision-language models (VLMs) exhibit limited performance on medical imaging benchmarks and rely heavily on manual prompt engineering, hindering scalability and clinical deployment. Method: We propose the first structured, automated prompt optimization framework tailored for medical VLMs—introducing DSPy to medical vision-language systems for the first time. Our framework enables end-to-end prompt automation without modifying model weights, abstracting prompt design into a modular pipeline and integrating four complementary optimization techniques. It supports five major medical imaging domains (radiology, gastroenterology, dermatology, pathology, and ophthalmology) and evaluates ten open-source VLMs. Contribution/Results: Experiments show a median relative performance gain of 53% over zero-shot baselines, with task-specific improvements up to 300–3400%. The framework supports privacy-preserving local deployment and is fully open-sourced.

Improving clinical image interpretation accuracy through automated methodsOptimizing vision-language models for medical imaging tasksReducing dependence on manual prompt engineering in healthcare

Measuring and Predicting Where and When Pathologists Focus their Visual Attention while Grading Whole Slide Images of Cancer

Aug 03, 2025
SC
Souradeep Chakraborty
🏛️ Stony Brook University | Northwell Health Laboratories | University of California San Francisco | University of Utah School of Medicine

This study addresses the problem of predicting pathologists’ spatiotemporal visual attention distributions while reviewing whole-slide images (WSIs) for cancer diagnosis. To model dynamic scanning trajectories, we propose a two-stage Transformer architecture: the first stage generates multi-scale attention heatmaps, and the second stage autoregressively predicts fixation sequences. We further introduce a semantics-preserving fixation extraction algorithm that jointly captures magnification level, spatial coordinates, and temporal dynamics. The model integrates digital microscope trajectory data with multi-scale histopathological features. Evaluated on 123 WSIs, it significantly outperforms random and baseline methods. This work presents the first end-to-end prediction framework for expert-level WSI scanning paths. It provides a quantifiable, interpretable attention assessment tool for pathology training and advances intelligent systems for diagnostic assistance and medical education.

Improve pathology training via attention predictionModel spatio-temporal focus during slide gradingPredict pathologists' visual attention on cancer slides

This work addresses the poor generalization of glioma subregion segmentation models and their limited responsiveness to precise clinical text-based corrections by proposing a lightweight, interactive segmentation framework built upon the frozen-weight 3D vision-language foundation model VoxTell. For the first time, the pretrained VoxTell is leveraged for text-instructed tumor segmentation editing: textual prompts encoding target, action, and spatial information are projected via a trainable adapter into the multi-scale decoder’s conditional pathway, enabling semantically controllable refinement without end-to-end retraining. On the BraTS-GLI test set, accurate instructions improved the Dice coefficient from 0.774 to 0.796, and cross-dataset experiments demonstrated statistically significant superiority over null or contradictory instructions (p<0.001).

glioma subregion segmentationinteractive correctionmedical image segmentation

Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.

binary segmentationevaluation metricsmetric decomposition

Latest Papers

What's happening recently
View more

This study addresses the unresolved question in LiDAR semantic scene completion of whether test-time iterative refinement or increasing model width yields better performance. Through rigorously controlled experiments with matched computational budgets, the authors compare single-pass prediction, width-scaled models, and a weight-sharing multi-grid iterative optimizer built upon a frozen backbone, evaluating robustness across controlled point cloud degradation modes—including coherent occlusion, independent sparsification, range-based attenuation, and clutter. Results demonstrate that iterative optimization provides substantial gains only under coherent missing regions (e.g., +0.911 mIoU on SemanticKITTI sequence 08, 95% CI [0.804, 1.040]), offers marginal improvement under sparse missingness (+0.300 points), and fails to handle clutter, at a cost of 10.74 ms and 0.75 GiB per frame. This work is the first to reveal the geometric dependency of iterative strategies under consistent training–testing conditions, offering empirical guidance for test-time compute allocation.

evidence geometryiterative vs one-shot predictionLiDAR scene completion

Prostate cancer exhibits subtle appearances on MRI and suffers from scarce annotations, making it challenging for existing automatic segmentation methods to balance accuracy and efficiency. This work proposes an interactive segmentation framework that, for the first time, integrates reinforcement learning with region growing. Leveraging point-based prompts, an uncertainty-guided exploration strategy, and an adaptive reward mechanism, the method achieves high-precision segmentation with minimal human intervention. Evaluated on the PROMIS and PI-CAI datasets, it outperforms the current state-of-the-art fully automatic approaches by 9.9% and 8.9%, respectively, attaining segmentation performance comparable to that of radiologists while reducing annotation time by an order of magnitude.

expert annotationimage-guided interventionmagnetic resonance imaging

This work addresses the limitations of the traditional Panoptic Quality (PQ) metric, which lacks a well-defined instance matching mechanism when IoU thresholds fall below 0.5, rendering it vulnerable to challenges such as fragmentation, ambiguous boundaries, and annotation noise. The authors formulate instance matching as a constrained bipartite graph assignment problem, decoupling match confidence from both prediction and ground truth sides. They systematically define four distinct matching strategies and introduce, for the first time, a vertex-centric framework that unifies the computation of true positives, false negatives, and false positives. This approach comprehensively characterizes the space of matching strategies under low-IoU conditions and naturally extends to part-aware panoptic segmentation evaluation—particularly beneficial for biomedical image analysis. The authors further release Panoptica, an open-source evaluation toolkit supporting multi-strategy and part-level assessment, demonstrating its efficacy across multiple case studies.

evaluation metricinstance matchingmatching strategy

This work addresses the longstanding disconnect between defect localization and structured reporting in industrial inspection, which has traditionally relied on manual intervention. The authors propose a decoupled three-stage pipeline: the Eyes module leverages YOLOv8-x-obb for high-precision oriented defect detection; the Bridge module maps detection outputs to structured prompts via parameter-free spatial encoding; and the Brain module employs a 4-bit quantized Qwen-2.5-1.5B model, fine-tuned with QLoRA and retrieval-augmented fine-tuning (RAFT), to generate standardized JSON reports. Evaluated on a small-scale synthetic dataset, the approach significantly outperforms generic end-to-end large models, achieving a BLEU-4 score of 0.41, a hallucination rate of only 4%, an expert rating of 8.6/10, and an inference speed of 47 tokens per second on a single T4 GPU—surpassing a 671B-parameter API baseline in overall performance.

defect localizationindustrial inspectionreport generation

This work addresses the limitations of existing surgical image segmentation methods, which are constrained by predefined categories, lack adaptive refinement capabilities, and do not support natural language interaction. To overcome these challenges, the authors propose IR-SIS, a novel interactive framework that leverages a fine-tuned SAM3 model to generate initial segmentations, integrates a vision-language model to assess segmentation quality and detect instruments, and employs an agent-based workflow to dynamically select refinement strategies. Crucially, IR-SIS introduces, for the first time, a human-in-the-loop mechanism that incorporates surgeons’ natural language feedback for iterative optimization. This approach transcends the static, closed-category paradigm of conventional methods, achieving state-of-the-art performance on both in-domain and out-of-distribution data from the EndoVis2017/2018 benchmarks, with surgeon interaction demonstrably enhancing segmentation accuracy.

adaptive refinementclinician interactionfoundation models

Hot Scholars

HH

Hui Huang

Chair Professor and CS Dean, Shenzhen University
GraphicsGeometryPointsShapes
RW

Ronggang Wang

Shenzhen Graduate School, Peking University
Immersive Video Coding and Processing
HJ

Hyun Jong Yang

Dept. of Electrical & Computer Engineering, Seoul National University
CommunicationsSignal ProcessingMachine Learning
CW

Chaoli Wang

Professor of Computer Science and Engineering, University of Notre Dame
Scientific VisualizationVisual AnalyticsVisualization
MK

Matthieu Komorowski

MD, PhD, Former Clinical Senior Lecturer at Imperial College London ; Visiting Scholar at MIT
AnesthesiaCritical CareMachine LearningSpace Medicine