keyframe selection

Selecting a sparse subset of frames that capture essential instance-level or anatomical information to enable open-vocabulary representations and reliably identify key anatomical planes in unconstrained video data.

keyframeselection

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection

Mar 05, 2025
DN
D. N. Kamtam
🏛️ Stanford University | Hotpot.ai

This study addresses the generalization bottleneck of anatomical tissue segmentation in surgical videos under few-shot, multi-organ, and cross-domain settings. We present the first adaptation of SAM 2 to surgical vision tasks, introducing a lightweight fine-tuning strategy that jointly optimizes the image encoder and mask decoder using only 50–400 samples per class. Segmentation is guided by point prompts (1–10 points) and evaluated via Weighted Mean Dice Coefficient (WMDC). Extensive multicenter validation across five datasets demonstrates a 17.9% relative improvement in WMDC over baselines (0.92 on validation sets); on test sets comprising 30 organ classes, our method outperforms prior state-of-the-art (SOTA) in 24 classes and achieves a 77.8% SOTA generalization rate for unseen organ categories. Our core contribution is the first efficient, video-aware SAM 2 fine-tuning paradigm for surgery—significantly enhancing segmentation accuracy and cross-domain robustness under zero-shot and few-shot conditions.

Achieving SOTA in segmentation for unseen organ classes.Evaluating dataset size impact on fine-tuning performance using WMDC.Fine-tuning SAM 2 for surgical anatomy segmentation in videos/images.

Interior Object Geometry via Fitted Frames

Jul 19, 2024
SM
Stephen M. Pizer
🏛️ University of North Carolina at Chapel Hill | Kitware, Inc.

To address the challenge of establishing stable inter-subject local correspondences in anatomical shape statistics—traditionally reliant on explicit registration—we propose a registration-free, boundary- and interior-consistent geometric modeling framework. Our method models target deformations as diffeomorphic transformations of an ellipsoid, embedded within a globally optimized skeleton-driven fitting scheme that simultaneously constructs a consistent coordinate system on both the object’s boundary and interior. We further introduce an evolutionary s-rep representation, the first to encode intrinsic geometric features directly in the fitted coordinate space, enabling robust point-wise correspondence across subjects without mesh alignment. The approach integrates differential-geometric deformation modeling, skeleton-guided fitting, and boundary-driven intrinsic coordinate generation. In hippocampal disease classification, it significantly outperforms two state-of-the-art methods, demonstrating superior discriminative power and statistical stability of the learned features.

Computing alignment-free geometric features for object populationsEnabling local correspondence in anatomic object representationsImproving classification performance via evolutionary s-rep model

Current medical multimodal large language models struggle to effectively interpret sparse, subtle, and context-dependent critical evidence embedded within full-length surgical videos encountered in real-world clinical settings. To address this gap, this work introduces the first long-context medical video understanding benchmark grounded in authentic clinical scenarios, comprising 759 hours of unsegmented, full-duration surgical videos and 1,253 evidence-based multiple-choice questions. The benchmark pioneers a joint evidence retrieval–reasoning evaluation paradigm with extremely sparse annotations—averaging only 0.166% of total video frames per question. Systematic evaluation reveals that state-of-the-art models achieve merely 41.1% accuracy, with performance failing to consistently improve as input frame count increases, exposing limitations in attentional focus amid redundant visual streams and deficiencies in multi-hop clinical reasoning. These findings underscore evidence localization and clinical interpretation as fundamental bottlenecks.

clinical reasoningin-the-wild benchmarklong-context medical video understanding

This work addresses the challenges in multi-view learning arising from severe dimensional imbalance, where high-dimensional views dominate, low-dimensional information is overlooked, and representation alignment becomes difficult. To tackle these issues, the authors propose AdaMuS, an adaptive sparse multi-view learning framework. AdaMuS employs view-specific encoders to map heterogeneous data into a unified space, incorporates a parameter-free pruning strategy to mitigate overfitting in low-dimensional views, and introduces a sparse fusion mechanism that adaptively removes redundancy and aligns multi-view representations. Furthermore, similarity graph-based self-supervised learning is leveraged to enhance model generalization. Extensive experiments on a synthetic dataset and seven real-world benchmarks demonstrate that AdaMuS achieves state-of-the-art performance in both classification and semantic segmentation tasks.

dimensional imbalancemulti-view learningoverfitting

Leveraging Foundation Models for Content-Based Medical Image Retrieval in Radiology

Mar 11, 2024
SD
Stefan Denner
🏛️ German Cancer Research Center (DKFZ) | Heidelberg University | Charité - Universitätsmedizin Berlin | Helmholtz Imaging | Heidelberg University Hospital | National Center for Tumor Diseases (NCT) Heidelberg

Existing radiological content-based image retrieval (CBIR) systems are typically disease-specific, exhibiting poor generalizability across pathologies and imaging modalities. To address this limitation, we propose the first general-purpose medical image retrieval paradigm leveraging weakly supervised vision foundation models—specifically ViT, CLIP, and DINOv2—without fine-tuning. Our approach supports cross-modal retrieval across four imaging modalities (e.g., X-ray, CT, MRI, ultrasound) and 161 distinct pathological categories. We rigorously evaluate it on a large-scale dataset of 1.6 million 2D radiological images, achieving a top-1 precision (P@1) of up to 0.594—comparable to state-of-the-art task-specific models. Crucially, we identify and empirically validate that retrieving pathology-related features is fundamentally more challenging than retrieving anatomical structures—a previously unreported insight. Our results demonstrate that foundation models exhibit strong generalization capability in large-scale, multi-disease CBIR, paving a novel pathway toward universal medical image retrieval systems.

Enhancing radiology diagnostics with foundation models for CBIREvaluating foundation models' performance on diverse radiological imagesOvercoming limitations of specialized CBIR systems with versatile models

Latest Papers

What's happening recently
View more

This work addresses the limited understanding of how frozen 3D medical vision encoders represent clinical findings, particularly which channels encode specific radiological concepts and their spatial locations. The authors propose Concept Channel Probing (CCP), a training-free method that, for the first time, reveals clinical concepts are encoded by only about ten sparse channels. The approach demonstrates cross-architecture generalizability across multiple 3D vision–language models. Experimental results show that CCP significantly outperforms CT-CHAT in clinical F1 score (0.549 vs. 0.184) and BLEU (0.483 vs. 0.373), while achieving a 22-fold reduction in inference latency. Further channel ablation studies confirm the specificity of the identified concept-channel mappings.

3D medical image interpretationclinical findingsconcept representation

Existing vision foundation models struggle to accurately localize functional keypoints—such as instrument tips and anchors—that are semantically tied to surgical actions like grasping or clamping in surgical videos. To address this, this work proposes a lightweight multi-frame network that leverages zero-shot point-prompt masks generated by SAM³ as dense structural priors. By fusing visual features with heatmap regression, the method achieves precise keypoint localization without requiring manual pixel-level annotations. The approach mitigates region bias inherent in direct mask supervision by using vision foundation model outputs as action-aware structural guidance. Evaluated across 7,867 surgical clips from five diverse datasets, the method achieves F1 scores of 72.4% for tip localization and 58.0% for anchor localization, demonstrating its effectiveness in heterogeneous surgical scenarios.

action-awarefunctional landmark localizationsparse localization

This study addresses the challenges of anatomical structure recognition in minimally invasive surgery, which are hindered by scarce annotated data and the bias of existing methods toward natural scenes. To overcome these limitations, the authors introduce ATLAS-120k, a large-scale semantic segmentation dataset comprising 120,000 video frames spanning 14 distinct surgical procedures. They propose the context-aware ATLAS model, which integrates foundation model embeddings with lightweight temporal reasoning, leveraging surgical type, procedural phase, and short-term visual memory to enhance both accuracy and temporal consistency. An innovative, scalable annotation pipeline combining expert labeling, automated propagation, and surgeon validation is also developed. The proposed approach achieves real-time inference while significantly improving segmentation performance, thereby laying a foundation for surgical navigation and clinical decision-support systems.

limited annotated dataminimally invasive surgerysemantic segmentation

This study addresses the challenge of accurately detecting small-scale anatomical structures—such as the inferior epigastric vessels—in laparoscopic inguinal hernia repair videos, where visual blur and intermittent visibility hinder reliable identification. To tackle this, the authors propose a Gaussian Spatial Prior (GSP) module that, for the first time, explicitly models spatial constraints among anatomical structures as learnable, compact Gaussian parameters. These priors are integrated into the self-attention mechanism of the DAB-DETR decoder and dynamically updated through iteratively refined reference points. Evaluated on a surgical video dataset of inguinal hernia repairs, the method significantly improves detection performance: compared to DAB-DETR and YOLOv26, it achieves relative gains of 33.5% and 53.9% in class-specific AP50, respectively, and enhances landmark detection accuracy by 6.0% (p=0.012), demonstrating the efficacy of anatomy-aware spatial priors for fine-grained object detection in surgical scenes.

anatomical object detectionintermittent visibilitysmall anatomical structures

This work addresses the high annotation cost and computational overhead of video instance segmentation, which typically relies on densely annotated frames. The authors propose a lightweight, end-to-end framework that achieves efficient training using only sparsely labeled frames. Key innovations include a Past-frames Feature Propagation (PFP) module for cross-frame feature propagation, lightweight frame-specific instance queries to model temporal evolution, and low-dimensional feature aggregation based on an image encoder. Evaluated on YouTube-VIS 2019/2021/2022 and OVIS benchmarks, the method attains near-full-supervision performance—within just a 0.4% AP gap—using only one-fifth of the annotated frames, significantly outperforming existing baselines and improving AP by over 1% under sparse supervision, thereby effectively bridging the accuracy gap between sparse and dense training paradigms.

Dense AnnotationsLabel EfficiencySparse Annotations

Hot Scholars

JP

Jun Peng

PhD, Soochow University, Australian National University
Photovoltaics
YS

Yusheng Su

AMD | Tsinghua University
Large Language ModelMachine LearningMLSys
EB

Emad Barsoum

AMD, Columbia University
Generative AIFoundation ModelsAgentic AIComputer Vision
AY

Alan Yuille

Professor of Cognitive Science and Computer Science, Johns Hopkins University
Computer VisionComputational Models of Mind and BrainMachine Learning