ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of excessively long visual token sequences when processing 3D CT scans with vision-language models, where existing compression methods struggle to balance anatomical fidelity and computational efficiency. The authors propose a training-free, plug-and-play compression approach that uniquely integrates organ-prior-guided token aggregation with region-wise centroid-based sinusoidal positional encoding. This method adaptively reduces token count while preserving critical anatomical information. Evaluated on the CT-RATE and Merlin datasets, it substantially outperforms baseline methods: under 64× visual context compression, it achieves a 50× reduction in KV cache memory and a 31× speedup in single-instance inference, while consistently enhancing performance in attribute prediction and medical text generation tasks.
📝 Abstract
A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed before a language model can consume it. Token compression is well studied in general vision, but little of it targets 3D CT specifically. A common baseline is grid average, which pools regular grid cells and can blend distinct anatomy, lesion, and air into one token. We present \textbf{ORCA} (ORgan-Centroid Aggregation), a token compressor for 3D CT. It merges adjacent tokens with organ guidance and adds a sinusoidal encoding of each region's centroid to preserve spatial layout. This preserves the anatomical information a downstream model needs. ORCA is training-free and plug-and-play, producing an adjustable token set without any model change or text query. We evaluate it across two datasets (CT-RATE and Merlin) and five encoders. The evaluation spans two task types: attribute prediction over five families (size, density, location, texture, and disease) and text generation (visual question answering and report generation). At matched token budgets, ORCA improves consistently over existing compression methods. It shrinks the visual context $64\times$ and its KV-cache $50\times$, and is $31\times$ faster to process each volume. Code released at https://github.com/renjie-liang/ORCA-3DCT.
Problem

Research questions and friction points this paper is trying to address.

3D CT
visual token compression
vision-language model
anatomical information preservation
token sequence length
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free
organ-guided aggregation
centroid encoding
3D CT token compression
plug-and-play
R
Renjie Liang
University of Florida, Gainesville, FL, USA
Z
Zijian Xu
University of Florida, Gainesville, FL, USA
J
Jinqian Pan
University of Florida, Gainesville, FL, USA
C
Chengkun Sun
University of Florida, Gainesville, FL, USA
Z
Zhengkang Fan
University of Florida, Gainesville, FL, USA
S
Shawn Li
University of Southern California, Los Angeles, CA, USA
Y
You Qin
National University of Singapore, Singapore
M
Mei Liu
University of Florida, Gainesville, FL, USA
Jie Xu
Jie Xu
Associate Professor, Electrical and Computer Engineering, University of Florida
Federated LearningMulti-Armed BanditsEdge ComputingWireless Networks