A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls
研究通过引入结合文本和音频的DualEvasion基准,解决财报电话会议中逃避检测问题,揭示现有模型在识别语音自信度上的不足。
研究通过引入结合文本和音频的DualEvasion基准,解决财报电话会议中逃避检测问题,揭示现有模型在识别语音自信度上的不足。
This work addresses the granularity mismatch between existing listwise explanation methods—which rely on isolated terms—and the semantic chunk representations leveraged by dense retrievers. To bridge this gap, the authors propose ChunkGroupSHAP, the first approach to incorporate semantic chunk clustering into Shapley value computation. By aligning attribution with the dense retriever’s representation granularity through cross-document semantic groupings, ChunkGroupSHAP enhances attribution consistency while preserving the listwise explanation framework. Experiments reveal that the optimal explanation unit varies with both the retriever and corpus: BM25 performs best with word-level units, dense models like E5 benefit from corpus-level groupings, and heterogeneous retrieval settings gain from query-local groupings. The method demonstrates consistent effectiveness across MS MARCO, FinanceBench, AILACaseDocs, and FinQA benchmarks.
This work addresses the scarcity of resources for detecting jailbreak attacks against vision-language models (VLMs) in the financial domain, where the integration of textual and visual modalities expands the attack surface. To bridge this gap, we introduce FENCE, the first bilingual (Chinese–English), finance-oriented multimodal jailbreak detection dataset. FENCE comprises adversarial samples generated from real-world financial image–text pairs, explicitly designed to contain harmful content and enable joint image–text threat modeling. Baseline detectors trained on FENCE achieve 99% accuracy on in-distribution data and demonstrate strong generalization on external benchmarks, underscoring the dataset’s effectiveness and practical utility in enhancing the security of financial AI systems.
Existing IR benchmarks inadequately capture the complex, domain-specific information needs inherent in banking operations, while constructing realistic, legally compliant evaluation datasets remains hindered by regulatory constraints and high annotation costs. To address this, we propose a systematic, large language model–based query generation framework that integrates single-document and multi-document joint sampling, augmented by a novel reasoning-enhanced answerability assessment mechanism—significantly improving alignment between generated queries and authentic business requirements. Our method explicitly supports modeling of complex multi-document retrieval scenarios. Leveraging it, we construct KoBankIR—the first Korean-language, banking-domain IR benchmark—comprising 815 high-quality, expert-validated queries. Empirical evaluation reveals substantial performance degradation of state-of-the-art retrieval models on multi-document tasks under KoBankIR, confirming its rigor and practical relevance. KoBankIR provides a reproducible, legally compliant, and high-fidelity evaluation foundation for financial-domain IR research.
研究通过引入结合文本和音频的DualEvasion基准,解决财报电话会议中逃避检测问题,揭示现有模型在识别语音自信度上的不足。
This work addresses the granularity mismatch between existing listwise explanation methods—which rely on isolated terms—and the semantic chunk representations leveraged by dense retrievers. To bridge this gap, the authors propose ChunkGroupSHAP, the first approach to incorporate semantic chunk clustering into Shapley value computation. By aligning attribution with the dense retriever’s representation granularity through cross-document semantic groupings, ChunkGroupSHAP enhances attribution consistency while preserving the listwise explanation framework. Experiments reveal that the optimal explanation unit varies with both the retriever and corpus: BM25 performs best with word-level units, dense models like E5 benefit from corpus-level groupings, and heterogeneous retrieval settings gain from query-local groupings. The method demonstrates consistent effectiveness across MS MARCO, FinanceBench, AILACaseDocs, and FinQA benchmarks.
This work addresses the scarcity of resources for detecting jailbreak attacks against vision-language models (VLMs) in the financial domain, where the integration of textual and visual modalities expands the attack surface. To bridge this gap, we introduce FENCE, the first bilingual (Chinese–English), finance-oriented multimodal jailbreak detection dataset. FENCE comprises adversarial samples generated from real-world financial image–text pairs, explicitly designed to contain harmful content and enable joint image–text threat modeling. Baseline detectors trained on FENCE achieve 99% accuracy on in-distribution data and demonstrate strong generalization on external benchmarks, underscoring the dataset’s effectiveness and practical utility in enhancing the security of financial AI systems.
Existing IR benchmarks inadequately capture the complex, domain-specific information needs inherent in banking operations, while constructing realistic, legally compliant evaluation datasets remains hindered by regulatory constraints and high annotation costs. To address this, we propose a systematic, large language model–based query generation framework that integrates single-document and multi-document joint sampling, augmented by a novel reasoning-enhanced answerability assessment mechanism—significantly improving alignment between generated queries and authentic business requirements. Our method explicitly supports modeling of complex multi-document retrieval scenarios. Leveraging it, we construct KoBankIR—the first Korean-language, banking-domain IR benchmark—comprising 815 high-quality, expert-validated queries. Empirical evaluation reveals substantial performance degradation of state-of-the-art retrieval models on multi-document tasks under KoBankIR, confirming its rigor and practical relevance. KoBankIR provides a reproducible, legally compliant, and high-fidelity evaluation foundation for financial-domain IR research.