Institution profile

Vietnamese - German University

Academic institutionasia · vn
Official website
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

Aug 13, 2026

This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.

0 citationsRead paper

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Aug 10, 2026

This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.

0 citationsRead paper
Recent publications

Latest Papers

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

Aug 13, 2026

This work addresses the challenge of fine-grained retrieval of anomalous pedestrian behaviors from large-scale image collections based on natural language descriptions. To tackle this problem, we propose a robust cross-modal retrieval framework that integrates heterogeneous vision-language embeddings through score alignment and iterative ensemble strategies to effectively fuse multi-model representations. Furthermore, we introduce a discrepancy-aware re-ranking mechanism to handle semantically ambiguous queries. The proposed approach significantly enhances the robustness and accuracy of cross-modal matching in complex scenarios, achieving state-of-the-art performance on the PAB benchmark with 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, thereby demonstrating its effectiveness.

0 citationsRead paper

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Aug 10, 2026

This work addresses the challenge of fine-grained matching in text-based pedestrian anomaly retrieval under synthetic-to-real (Sim2Real) scenarios. To this end, the authors propose an anchor-constrained coarse-to-fine retrieval framework that leverages multi-facet semantic decomposition and calibrated fusion. The approach integrates a heterogeneous vision-language retriever, a Qwen3-based reranker, and an anomaly-aware cloze-style verification module, complemented by an uncertainty-gated consensus mechanism operating over a small candidate pool to enable efficient fine-grained semantic reasoning. Innovatively, semantic facets serve as anchor constraints to jointly optimize recall and computational efficiency. Evaluated on the PAB benchmark, the method achieves 95.41% mAP@10, 94.44% R@1, and 99.09% R@5, significantly outperforming existing single-backbone models.

0 citationsRead paper