Institution profile

Guangxi University

Academic institutionasia · cn
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

Aug 02, 2026

This work addresses the challenge in zero-shot image captioning where synthetic training data generated by text-to-image models often suffers from fine-grained entity misalignment—such as missing objects or mislocalized attributes—leading to distorted supervision signals. To mitigate this, the authors propose ReCap, a framework that explicitly detects image entities and guides caption rewriting to achieve fine-grained image-text alignment. ReCap further incorporates an adaptive dynamic weighting strategy to downweight unreliable synthetic samples during training. By shifting data refinement from implicit global matching to explicit entity-level realignment, the method introduces a plug-and-play mechanism for fine-grained correction. Experiments demonstrate that ReCap achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks, significantly improving caption consistency and supervision fidelity.

0 citationsRead paper
Recent publications

Latest Papers

Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

Aug 02, 2026

This work addresses the challenge in zero-shot image captioning where synthetic training data generated by text-to-image models often suffers from fine-grained entity misalignment—such as missing objects or mislocalized attributes—leading to distorted supervision signals. To mitigate this, the authors propose ReCap, a framework that explicitly detects image entities and guides caption rewriting to achieve fine-grained image-text alignment. ReCap further incorporates an adaptive dynamic weighting strategy to downweight unreliable synthetic samples during training. By shifting data refinement from implicit global matching to explicit entity-level realignment, the method introduces a plug-and-play mechanism for fine-grained correction. Experiments demonstrate that ReCap achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks, significantly improving caption consistency and supervision fidelity.

0 citationsRead paper