Institution profile

Huaqiao University

Academic institutionasia · cn
Official website
Research library47linked papers
Opportunities0open roles
Selected work

Representative Papers

Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction

Oct 03, 2026

This study addresses the vulnerability of speech synthesis watermarking to model-driven reconstruction attacks, which undermines reliable source tracing. To this end, we propose Thrive, a framework for multi-bit generative watermarking in autoregressive text-to-speech (TTS) systems. Thrive introduces a novel adaptive multi-bit mechanism that eliminates reliance on a single carrier and supports both discrete token and continuous representation paradigms. Specifically, it synchronously injects intermediate representations via a Rise module and employs a Care module that integrates waveform and spectrogram experts for bit-level reliability selection. Experimental results demonstrate that Thrive maintains high synthesis fidelity while achieving an 87.6% watermark recovery accuracy under reconstruction attacks, thereby enabling identity tracing across tens of thousands of users.

0 citationsRead paper

Geometry-based Schrödinger Bridges for Trustworthy Multimodal Fusion

May 29, 2026

This work addresses a critical limitation in existing trustworthy multimodal fusion methods, which rely on model prediction confidence to assess input quality and consequently fail when models are confidently incorrect. To overcome this circular dependency between confidence and correctness, the authors propose a geometry-inspired reliability criterion. Specifically, they model the transport path from an input to a reliable region in latent space using a Rectified Flow–based diffusion Schrödinger bridge, and define a calibration score as the squared norm of the initial transport velocity. This score provides a model-confidence–agnostic measure of input reliability. Empirical results demonstrate that the proposed approach significantly outperforms current baselines under challenging conditions involving strong sensor noise and semantic conflicts, thereby substantially enhancing the robustness of multimodal fusion.

0 citationsRead paper

How Far Has AI Come in Liver Fibrosis Staging? A Large-Scale Real-World Dataset and Benchmark

May 25, 2026

This study addresses the lack of systematic evaluation of AI performance in staging liver fibrosis within real-world, multicenter, and heterogeneous clinical settings. To this end, we constructed LiFS—the first large-scale multicenter dataset comprising complete gadoxetic acid–enhanced multiphase MRI sequences paired with histopathological reference standards—and leveraged the MICCAI 2025 CARE-Liver Challenge to systematically benchmark nine AI approaches. Through strategies including multiseries registration, multimodal fusion, and diverse backbone architectures with varying input dimensionalities, the top-performing model achieved diagnostic accuracy comparable to that of experienced radiologists and significantly outperformed junior readers. Our findings highlight inter-center heterogeneity, label imbalance, and variability in contrast-enhancement protocols as key challenges, offering critical benchmarks and insights for future clinical deployment of AI in liver fibrosis assessment.

0 citationsRead paper

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

May 21, 2026

This work addresses the challenge of cross-modal noise propagation caused by unidirectional or asynchronous manipulations in audio-visual deepfakes. To this end, we propose IaMSB, a novel framework that, for the first time, introduces the Schrödinger bridge into forgery localization. Our method employs a lightweight coarse bridge to screen candidate intervals and estimate cross-modal consistency, followed by a refined bridge that performs asymmetric step optimization and bottlenecked cross-modal interaction to achieve high-precision interval-level temporal localization. IaMSB unifies cross-modal consistency modeling, informative segment selection, and bridge-step scheduling within a single architecture. Extensive experiments demonstrate significant performance gains under strict IoU thresholds across multiple benchmarks, with AP@0.95 improvements of 3%–10%, particularly excelling in scenarios involving unidirectional forgeries.

0 citationsRead paper

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

Apr 16, 2026

Existing video understanding methods struggle to model objects undergoing semantically significant changes over time and lack explicit reasoning over critical visual evidence. This work proposes a search-guided, progressive object grounding framework that incrementally anchors task-relevant visual regions through a reinforcement learning–driven search controller coupled with a novel formatting reward mechanism. By explicitly incentivizing the model to focus on authentic visual evidence, the approach constructs spatially grounded, multi-step reasoning trajectories. Notably, it introduces formatting rewards into visual reasoning for the first time. The method consistently achieves performance gains across multiple benchmarks—including NExTQA, Video-Holmes, CG-Bench Reasoning, and VRBench—demonstrating superior accuracy, interpretability, robustness, and cross-domain generalization.

0 citationsRead paper
Recent publications

Latest Papers

Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction

Oct 03, 2026

This study addresses the vulnerability of speech synthesis watermarking to model-driven reconstruction attacks, which undermines reliable source tracing. To this end, we propose Thrive, a framework for multi-bit generative watermarking in autoregressive text-to-speech (TTS) systems. Thrive introduces a novel adaptive multi-bit mechanism that eliminates reliance on a single carrier and supports both discrete token and continuous representation paradigms. Specifically, it synchronously injects intermediate representations via a Rise module and employs a Care module that integrates waveform and spectrogram experts for bit-level reliability selection. Experimental results demonstrate that Thrive maintains high synthesis fidelity while achieving an 87.6% watermark recovery accuracy under reconstruction attacks, thereby enabling identity tracing across tens of thousands of users.

0 citationsRead paper

Geometry-based Schrödinger Bridges for Trustworthy Multimodal Fusion

May 29, 2026

This work addresses a critical limitation in existing trustworthy multimodal fusion methods, which rely on model prediction confidence to assess input quality and consequently fail when models are confidently incorrect. To overcome this circular dependency between confidence and correctness, the authors propose a geometry-inspired reliability criterion. Specifically, they model the transport path from an input to a reliable region in latent space using a Rectified Flow–based diffusion Schrödinger bridge, and define a calibration score as the squared norm of the initial transport velocity. This score provides a model-confidence–agnostic measure of input reliability. Empirical results demonstrate that the proposed approach significantly outperforms current baselines under challenging conditions involving strong sensor noise and semantic conflicts, thereby substantially enhancing the robustness of multimodal fusion.

0 citationsRead paper

How Far Has AI Come in Liver Fibrosis Staging? A Large-Scale Real-World Dataset and Benchmark

May 25, 2026

This study addresses the lack of systematic evaluation of AI performance in staging liver fibrosis within real-world, multicenter, and heterogeneous clinical settings. To this end, we constructed LiFS—the first large-scale multicenter dataset comprising complete gadoxetic acid–enhanced multiphase MRI sequences paired with histopathological reference standards—and leveraged the MICCAI 2025 CARE-Liver Challenge to systematically benchmark nine AI approaches. Through strategies including multiseries registration, multimodal fusion, and diverse backbone architectures with varying input dimensionalities, the top-performing model achieved diagnostic accuracy comparable to that of experienced radiologists and significantly outperformed junior readers. Our findings highlight inter-center heterogeneity, label imbalance, and variability in contrast-enhancement protocols as key challenges, offering critical benchmarks and insights for future clinical deployment of AI in liver fibrosis assessment.

0 citationsRead paper

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

May 21, 2026

This work addresses the challenge of cross-modal noise propagation caused by unidirectional or asynchronous manipulations in audio-visual deepfakes. To this end, we propose IaMSB, a novel framework that, for the first time, introduces the Schrödinger bridge into forgery localization. Our method employs a lightweight coarse bridge to screen candidate intervals and estimate cross-modal consistency, followed by a refined bridge that performs asymmetric step optimization and bottlenecked cross-modal interaction to achieve high-precision interval-level temporal localization. IaMSB unifies cross-modal consistency modeling, informative segment selection, and bridge-step scheduling within a single architecture. Extensive experiments demonstrate significant performance gains under strict IoU thresholds across multiple benchmarks, with AP@0.95 improvements of 3%–10%, particularly excelling in scenarios involving unidirectional forgeries.

0 citationsRead paper

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

Apr 16, 2026

Existing video understanding methods struggle to model objects undergoing semantically significant changes over time and lack explicit reasoning over critical visual evidence. This work proposes a search-guided, progressive object grounding framework that incrementally anchors task-relevant visual regions through a reinforcement learning–driven search controller coupled with a novel formatting reward mechanism. By explicitly incentivizing the model to focus on authentic visual evidence, the approach constructs spatially grounded, multi-step reasoning trajectories. Notably, it introduces formatting rewards into visual reasoning for the first time. The method consistently achieves performance gains across multiple benchmarks—including NExTQA, Video-Holmes, CG-Bench Reasoning, and VRBench—demonstrating superior accuracy, interpretability, robustness, and cross-domain generalization.

0 citationsRead paper