๐ค AI Summary
This work addresses the challenge in zero-shot compositional image retrieval of simultaneously preserving visual continuity and accurately reflecting textual semantic modifications, a dilemma that often leads existing methods to suffer from perceptual myopia or logical drift. To overcome this, the authors propose an end-to-end hierarchical Perception-Reasoning Framework (PDF), which introducesโ for the first timeโan experience-based self-evolution mechanism and a test-time scaling law (TTS). Leveraging a hierarchical multi-agent architecture, an intent-aware routing manager, and a training-free strategy for distilling reasoning capabilities, PDF enables fine-grained collaborative inference at test time. The method achieves state-of-the-art performance across CIRR, CIRCO, and FashionIQ benchmarks, significantly enhancing both model effectiveness and scalability.
๐ Abstract
Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the semantic variables specified in the modification text, which constitutes the core challenge of the task. Existing methods often suffer from Perception Myopia in a single space, or fall into Logic Drift in iterative collaboration due to the perception ceiling of the underlying retriever. To address this issue, we propose a one-stop hierarchical Perception-to-Deliberation Framework (PDF), which, to the best of our knowledge, is the first to introduce experience self-evolution and Test-Time Scaling Law (TTS) into ZS-CIR. Relying on a hierarchical multi-agent architecture, PDF first utilizes an Intent Routing Manager to dynamically dispatch multi-view Worker perception signals based on modification intents to construct a high-recall candidate pool. Subsequently, the Decision Manager combines a Training-free Reasoning Policy Distillation mechanism with a Tournament-style TTS strategy to achieve self-evolving fine-grained reasoning, yielding the final retrieval results. Experimental results demonstrate that PDF achieves SOTA performance on three benchmark datasets: CIRR, CIRCO, and FashionIQ. This study indicates that experience-driven self-evolution and TTS represent a highly promising and scalable path for achieving zero-shot fine-grained multimedia retrieval. The code will be made publicly available upon acceptance.