🤖 AI Summary
This work addresses the challenge of efficiently identifying optimal configurations in 3D-CT vision-language models, where combining frozen image encoders with token compression schemes typically requires costly fine-tuning. To circumvent this, the authors propose a low-cost probing method that predicts downstream fine-tuning performance directly from cached embeddings of frozen encoders. They introduce an image-guided probing benchmark to evaluate the ranking efficacy of various (encoder × compression) combinations and incorporate two novel validation mechanisms—scale-sanity and probe-separability—to ensure clinically relevant attributes are decodable and representation scales remain reasonable. Experimental results demonstrate that the probe-based rankings correlate strongly with full fine-tuning outcomes (Spearman’s ρ ≈ 0.95), enabling candidate configuration screening within minutes and substantially reducing computational overhead.
📝 Abstract
Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their tokens, and several token budgets, and the combinations grow fast. Comparing them the usual way means fine-tuning a large language model (LLM) on each combination, and running the whole sweep this way needs far more compute than most groups can spend. We ask whether a cheap probe on the encoder's cached embeddings can stand in for that comparison. We build an image-grounded probing benchmark over (encoder $\times$ compression) cells, with clinical attribute families and two validation gates, scale-sanity and probe-separability, that keep each attribute well-scaled and decodable. These gates are the main methodological contribution. On this benchmark we compare a range of read-out heads, and in a preliminary study we pair each probe with its matched downstream task. The early signal is encouraging: the cheap probe orders the candidates in close agreement with expensive fine-tuning, at about $r\approx0.95$ on the cells measured so far. We read this as an ordinal claim, a ranking predictor rather than an exact estimate, and we are explicit about where it stays preliminary. If it holds up, encoder and compression choices can be screened in minutes with frozen-token probes, with full training spent only on the finalists.