Institution profile

LMMs-Lab

Research institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

Sep 29, 2026

This study addresses the fine-grained perception challenges in high-resolution scenarios, namely visual token redundancy and context loss caused by local cropping. We propose a plug-and-play, lightweight evidence-adaptive framework that pioneers the use of human visual search trajectories to supervise evidence density, thereby guiding region re-reading and dynamic token allocation. Furthermore, a sparse coordinate bridging network is designed to integrate local features with the global scene. The proposed method can be deployed without retraining the host backbone network, ensuring strong compatibility. Experimental results demonstrate that our approach significantly improves average fine-grained accuracy across nine backbone models, outperforming purely global methods under identical token budgets.

0 citationsRead paper

HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

Sep 29, 2026

This study addresses the issue of erroneous reasoning paths in multi-turn visual search agents caused by reliance on outcome-only rewards, proposing a pioneering process reinforcement learning framework anchored in human search behavior to achieve a paradigm shift from outcome supervision to process supervision. Methodologically, this work integrates a manual annotation platform, a task-adaptive weight scorer, and a trajectory distillation-based process reward model to supervise search trajectories with fine-grained data, thereby optimizing high-resolution image question-answering performance. Experimental results demonstrate that the proposed framework significantly outperforms conventional outcome-oriented reinforcement learning approaches. Furthermore, applying process supervision during early training stages improves subsequent scaling efficiency by 6.7 times.

0 citationsRead paper
Recent publications

Latest Papers

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

Sep 29, 2026

This study addresses the fine-grained perception challenges in high-resolution scenarios, namely visual token redundancy and context loss caused by local cropping. We propose a plug-and-play, lightweight evidence-adaptive framework that pioneers the use of human visual search trajectories to supervise evidence density, thereby guiding region re-reading and dynamic token allocation. Furthermore, a sparse coordinate bridging network is designed to integrate local features with the global scene. The proposed method can be deployed without retraining the host backbone network, ensuring strong compatibility. Experimental results demonstrate that our approach significantly improves average fine-grained accuracy across nine backbone models, outperforming purely global methods under identical token budgets.

0 citationsRead paper

HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

Sep 29, 2026

This study addresses the issue of erroneous reasoning paths in multi-turn visual search agents caused by reliance on outcome-only rewards, proposing a pioneering process reinforcement learning framework anchored in human search behavior to achieve a paradigm shift from outcome supervision to process supervision. Methodologically, this work integrates a manual annotation platform, a task-adaptive weight scorer, and a trajectory distillation-based process reward model to supervise search trajectories with fine-grained data, thereby optimizing high-resolution image question-answering performance. Experimental results demonstrate that the proposed framework significantly outperforms conventional outcome-oriented reinforcement learning approaches. Furthermore, applying process supervision during early training stages improves subsequent scaling efficiency by 6.7 times.

0 citationsRead paper