🤖 AI Summary
This study addresses the limitations of zero-shot Chinese character recognition, where global matching overlooks the spatial layout of radicals and fine-grained ranking fails to capture subtle differences. To this end, this work proposes a global-to-local two-stage framework. In the retrieval stage, building upon a CLIP architecture with Ideographic Description Sequence (IDS) encoding, spatially aware prototypes are constructed by incorporating explicit tree positions and radical-level geometric priors, enabling high-recall candidate retrieval. In the re-ranking stage, a margin-gated radical verification mechanism is designed to enhance local discriminability through instance-query matching, thereby achieving precise ranking. The proposed method attains a Top-1 accuracy of 83.06% on the ICDAR2013 benchmark, establishing state-of-the-art performance. Comprehensive ablation studies further validate the effectiveness of each individual module within the framework.
📝 Abstract
Zero-shot Chinese character recognition (ZS-CCR) aims to recognize characters whose categories are never observed during training, and typically relies on the compositional structure shared between seen and unseen characters. Recent CLIP-style methods represent this structure with the Ideographic Description Sequence (IDS) and align it with glyph images in a shared embedding space. However, they rely on a single global image--IDS similarity that discards the spatial layout of radicals and, being learned only implicitly from seen classes, generalizes poorly to unseen ones; moreover, global matching often retrieves the correct character within the top candidates yet fails to rank it first when characters differ only in subtle local radicals. To address these issues, we propose a global-to-local two-stage framework. In the first stage, STG-CLIP augments the IDS with explicit tree-position and radical-level geometric priors, yielding a spatial-aware prototype that provides a consistent spatial description across seen and unseen categories for high-recall global retrieval. In the second stage, the Radical Verification Module (RVM) uses the radical instances of each retrieved candidate as queries to verify whether the corresponding radicals can be matched to spatially compatible regions in the input glyph. A margin-based gating rule activates the RVM only when the leading global candidates receive similar similarity scores. Experiments on the ICDAR2013 benchmark demonstrate that our method achieves state-of-the-art performance under the character-level zero-shot setting, obtaining 83.06% top-1 accuracy with 2,755 seen classes. Ablation studies further show that the explicit geometric priors and radical-level verification provide complementary improvements.