🤖 AI Summary
This study addresses the prohibitive latency and energy consumption associated with whole-slide inference of high-resolution images on satellite edge devices. To this end, it introduces Rift, a system that formalizes the concept of answer-invariant token redundancy for the first time. Leveraging vision-language models, Rift achieves dynamic computational resource allocation through query-conditioned patch pruning and further optimizes the inference pipeline via elastic prefilling. The proposed system has been deployed on the Jetson AGX Orin platform. Experimental results demonstrate that, compared to baseline methods, Rift improves accuracy from 45% to 73% while simultaneously reducing inference latency by 69% and energy consumption by 78%, thereby achieving synergistic enhancements in both performance and efficiency for edge computing scenarios.
📝 Abstract
Onboard vision-language models could enable satellites to answer queries directly, but exhaustive tiled inference over high-resolution imagery is slow and energy-intensive. We identify answer-invariant token redundancy (AITR): image tiles and vision tokens that can be removed without changing the final answer. We present Rift, a two-stage system that performs query-conditioned tile pruning followed by elastic prefill to reduce token budget. We evaluate it on LLaVA-1.5 7B running on Jetson AGX Orin. Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.