Institution profile

E.SUN Financial Holding

Industry researchasia · tw
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

From Global Alignment to Local Grounding: Zero-Shot Chinese Character Recognition with Radical Verification

Oct 07, 2026

This study addresses the limitations of zero-shot Chinese character recognition, where global matching overlooks the spatial layout of radicals and fine-grained ranking fails to capture subtle differences. To this end, this work proposes a global-to-local two-stage framework. In the retrieval stage, building upon a CLIP architecture with Ideographic Description Sequence (IDS) encoding, spatially aware prototypes are constructed by incorporating explicit tree positions and radical-level geometric priors, enabling high-recall candidate retrieval. In the re-ranking stage, a margin-gated radical verification mechanism is designed to enhance local discriminability through instance-query matching, thereby achieving precise ranking. The proposed method attains a Top-1 accuracy of 83.06% on the ICDAR2013 benchmark, establishing state-of-the-art performance. Comprehensive ablation studies further validate the effectiveness of each individual module within the framework.

0 citationsRead paper

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

Jun 08, 2026

Existing automatic evaluation methods for text-to-music generation systems struggle to optimize ranking metrics and exhibit weak cross-modal consistency. This work proposes DeRA-MOS, a novel framework that decouples listwise ranking from modality alignment objectives for the first time. It employs a batch-aware listwise ranking loss to optimize the ranking performance of musical impressions and integrates a score-anchored modality alignment loss to enhance semantic consistency between text and music. By explicitly addressing pointwise training bias and modality drift, the proposed approach significantly improves Spearman rank correlation on the MusicEval benchmark, establishing a new paradigm for large-scale evaluation of text-to-music generation systems.

0 citationsRead paper

Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations

Mar 17, 2026

This work addresses the vulnerability of existing automatic audio quality assessment models to data scarcity, which often leads them to learn spurious acoustic correlations tied to specific datasets rather than genuine perceptual quality characteristics. To mitigate this, the authors propose a disentanglement framework based on Domain-Adversarial Training (DAT), systematically exploring diverse domain definition strategies—from explicit metadata to implicit clustering—to separate quality-relevant features from confounding factors. A key insight is that no universally optimal domain partitioning exists; instead, the choice of strategy should be adaptively tailored to different Mean Opinion Score (MOS) dimensions. Experimental results demonstrate that the proposed approach significantly improves correlation with human ratings and exhibits superior generalization to unseen audio generation scenarios.

0 citationsRead paper

HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation

Feb 16, 2026

This work proposes HyperRAG, a novel retrieval-augmented generation (RAG) framework that addresses the limitations of traditional binary knowledge graph–based approaches in multi-hop question answering—namely, rigid retrieval, high computational cost, and insufficient relational expressiveness. HyperRAG is the first to incorporate n-ary hypergraphs into RAG, modeling high-order relational facts through hypergraph structures. It integrates structural-semantic joint reasoning with parametric memory from large language models, featuring a HyperRetriever module that adaptively constructs multi-hop reasoning paths and a HyperMemory mechanism that dynamically guides path expansion. Evaluated on benchmarks including WikiTopics and HotpotQA, HyperRAG significantly outperforms existing methods, achieving an average improvement of 2.95% in MRR and 1.23% in Hits@10, while also demonstrating enhanced interpretability and cross-domain generalization capabilities.

0 citationsRead paper
Recent publications

Latest Papers

From Global Alignment to Local Grounding: Zero-Shot Chinese Character Recognition with Radical Verification

Oct 07, 2026

This study addresses the limitations of zero-shot Chinese character recognition, where global matching overlooks the spatial layout of radicals and fine-grained ranking fails to capture subtle differences. To this end, this work proposes a global-to-local two-stage framework. In the retrieval stage, building upon a CLIP architecture with Ideographic Description Sequence (IDS) encoding, spatially aware prototypes are constructed by incorporating explicit tree positions and radical-level geometric priors, enabling high-recall candidate retrieval. In the re-ranking stage, a margin-gated radical verification mechanism is designed to enhance local discriminability through instance-query matching, thereby achieving precise ranking. The proposed method attains a Top-1 accuracy of 83.06% on the ICDAR2013 benchmark, establishing state-of-the-art performance. Comprehensive ablation studies further validate the effectiveness of each individual module within the framework.

0 citationsRead paper

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

Jun 08, 2026

Existing automatic evaluation methods for text-to-music generation systems struggle to optimize ranking metrics and exhibit weak cross-modal consistency. This work proposes DeRA-MOS, a novel framework that decouples listwise ranking from modality alignment objectives for the first time. It employs a batch-aware listwise ranking loss to optimize the ranking performance of musical impressions and integrates a score-anchored modality alignment loss to enhance semantic consistency between text and music. By explicitly addressing pointwise training bias and modality drift, the proposed approach significantly improves Spearman rank correlation on the MusicEval benchmark, establishing a new paradigm for large-scale evaluation of text-to-music generation systems.

0 citationsRead paper

Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations

Mar 17, 2026

This work addresses the vulnerability of existing automatic audio quality assessment models to data scarcity, which often leads them to learn spurious acoustic correlations tied to specific datasets rather than genuine perceptual quality characteristics. To mitigate this, the authors propose a disentanglement framework based on Domain-Adversarial Training (DAT), systematically exploring diverse domain definition strategies—from explicit metadata to implicit clustering—to separate quality-relevant features from confounding factors. A key insight is that no universally optimal domain partitioning exists; instead, the choice of strategy should be adaptively tailored to different Mean Opinion Score (MOS) dimensions. Experimental results demonstrate that the proposed approach significantly improves correlation with human ratings and exhibits superior generalization to unseen audio generation scenarios.

0 citationsRead paper

HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation

Feb 16, 2026

This work proposes HyperRAG, a novel retrieval-augmented generation (RAG) framework that addresses the limitations of traditional binary knowledge graph–based approaches in multi-hop question answering—namely, rigid retrieval, high computational cost, and insufficient relational expressiveness. HyperRAG is the first to incorporate n-ary hypergraphs into RAG, modeling high-order relational facts through hypergraph structures. It integrates structural-semantic joint reasoning with parametric memory from large language models, featuring a HyperRetriever module that adaptively constructs multi-hop reasoning paths and a HyperMemory mechanism that dynamically guides path expansion. Evaluated on benchmarks including WikiTopics and HotpotQA, HyperRAG significantly outperforms existing methods, achieving an average improvement of 2.95% in MRR and 1.23% in Hits@10, while also demonstrating enhanced interpretability and cross-domain generalization capabilities.

0 citationsRead paper