Institution profile

Université de Reims Champagne-Ardenne

Academic institutioneurope · fr
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing SubjectiveNLP Tasks

Apr 18, 2026

This work addresses the challenge of annotation disagreement in subjective NLP tasks, where ambiguity in labeling criteria or overlapping category boundaries often leads to inconsistent judgments. The authors propose a pattern-level diagnostic framework that introduces an interpretable, criterion-level auditing mechanism prior to label aggregation. By collecting fine-grained evaluations from multiple annotators on individual labeling guidelines, the method systematically identifies two failure modes: unstable annotation standards and systematic category overlap. This approach provides the first structured attribution of disagreement early in the annotation pipeline, offering empirical grounding for refining annotation schemes. Evaluated on a commercial document task involving persuasion-value extraction, the analysis reveals that disagreements concentrate around a few unstable criteria, with nearly half of the sentences activating multiple categories. The diagnostic outcomes show strong alignment with domain expert assessments and effectively inform revisions to the annotation guidelines.

0 citationsRead paper

Unsupervised Anomaly Detection in NSL-KDD Using $β$-VAE: A Latent Space and Reconstruction Error Approach

Feb 23, 2026

This work addresses the challenge of anomaly detection in unlabeled network traffic within IT/OT convergence scenarios by proposing an unsupervised approach based on β-variational autoencoders (β-VAE). It presents the first systematic comparison between two detection mechanisms: latent space distance and reconstruction error. Experimental evaluation on the NSL-KDD dataset demonstrates that measuring the distance from test samples to the training data distribution in the latent space significantly outperforms conventional reconstruction error–based methods, yielding notably higher detection accuracy under fully unsupervised conditions. The findings highlight the critical advantage of leveraging latent space structure for unsupervised network anomaly detection and offer a novel direction for future research in this domain.

0 citationsRead paper

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

Oct 20, 2025

Conventional evaluation relies heavily on scalar accuracy metrics, failing to characterize how models internally represent ambiguous samples—especially those exhibiting substantial human annotation disagreement. Method: This work pioneers the application of topological data analysis (TDA), specifically the Mapper algorithm, to the fine-tuned embedding space of RoBERTa-Large on the MD-Offense dataset, enabling systematic geometric characterization of ambiguity encoding. Unlike linear (e.g., PCA) or locally preserving (e.g., UMAP) dimensionality reduction methods, Mapper captures non-convex, modular decision regions in high-dimensional embeddings. Contribution/Results: We identify that over 98% of Mapper-generated connected components achieve ≥90% prediction purity and localize critical failure modes—including boundary collapse and overconfident clusters. Furthermore, we introduce the first topology-driven, connectivity-based quantitative metric, uncovering an implicit tension between “structurally high-confidence” representations and “label-level low-certainty.” This establishes a novel paradigm for probing ambiguity modeling mechanisms in large language models.

0 citationsRead paper

The information flow among Green Bonds exchange traded funds

Sep 23, 2025

This study investigates the direction and magnitude of information spillovers among 13 green bond ETFs across U.S., Canadian, and European markets during 2021–2022, aiming to uncover the intermarket linkage structure and dominance mechanisms in global green finance. Methodologically, it pioneers the systematic application of Transfer Entropy (TE) and Effective Transfer Entropy (ETE) to quantify nonlinear, directional, and causal information flows—overcoming limitations of conventional linear approaches. Results reveal that the U.S.-listed FLMB ETF acts as the strongest information source, the Europe-listed KLMH.F serves as the largest information sink, and HGGB exhibits significant cross-market propagation power across all three regions; overall, the U.S. market dominates the information network. This work provides novel empirical evidence and a methodological framework for understanding financial integration and systemic risk transmission in green bond markets.

0 citationsRead paper
Recent publications

Latest Papers

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing SubjectiveNLP Tasks

Apr 18, 2026

This work addresses the challenge of annotation disagreement in subjective NLP tasks, where ambiguity in labeling criteria or overlapping category boundaries often leads to inconsistent judgments. The authors propose a pattern-level diagnostic framework that introduces an interpretable, criterion-level auditing mechanism prior to label aggregation. By collecting fine-grained evaluations from multiple annotators on individual labeling guidelines, the method systematically identifies two failure modes: unstable annotation standards and systematic category overlap. This approach provides the first structured attribution of disagreement early in the annotation pipeline, offering empirical grounding for refining annotation schemes. Evaluated on a commercial document task involving persuasion-value extraction, the analysis reveals that disagreements concentrate around a few unstable criteria, with nearly half of the sentences activating multiple categories. The diagnostic outcomes show strong alignment with domain expert assessments and effectively inform revisions to the annotation guidelines.

0 citationsRead paper

Unsupervised Anomaly Detection in NSL-KDD Using $β$-VAE: A Latent Space and Reconstruction Error Approach

Feb 23, 2026

This work addresses the challenge of anomaly detection in unlabeled network traffic within IT/OT convergence scenarios by proposing an unsupervised approach based on β-variational autoencoders (β-VAE). It presents the first systematic comparison between two detection mechanisms: latent space distance and reconstruction error. Experimental evaluation on the NSL-KDD dataset demonstrates that measuring the distance from test samples to the training data distribution in the latent space significantly outperforms conventional reconstruction error–based methods, yielding notably higher detection accuracy under fully unsupervised conditions. The findings highlight the critical advantage of leveraging latent space structure for unsupervised network anomaly detection and offer a novel direction for future research in this domain.

0 citationsRead paper

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

Oct 20, 2025

Conventional evaluation relies heavily on scalar accuracy metrics, failing to characterize how models internally represent ambiguous samples—especially those exhibiting substantial human annotation disagreement. Method: This work pioneers the application of topological data analysis (TDA), specifically the Mapper algorithm, to the fine-tuned embedding space of RoBERTa-Large on the MD-Offense dataset, enabling systematic geometric characterization of ambiguity encoding. Unlike linear (e.g., PCA) or locally preserving (e.g., UMAP) dimensionality reduction methods, Mapper captures non-convex, modular decision regions in high-dimensional embeddings. Contribution/Results: We identify that over 98% of Mapper-generated connected components achieve ≥90% prediction purity and localize critical failure modes—including boundary collapse and overconfident clusters. Furthermore, we introduce the first topology-driven, connectivity-based quantitative metric, uncovering an implicit tension between “structurally high-confidence” representations and “label-level low-certainty.” This establishes a novel paradigm for probing ambiguity modeling mechanisms in large language models.

0 citationsRead paper

The information flow among Green Bonds exchange traded funds

Sep 23, 2025

This study investigates the direction and magnitude of information spillovers among 13 green bond ETFs across U.S., Canadian, and European markets during 2021–2022, aiming to uncover the intermarket linkage structure and dominance mechanisms in global green finance. Methodologically, it pioneers the systematic application of Transfer Entropy (TE) and Effective Transfer Entropy (ETE) to quantify nonlinear, directional, and causal information flows—overcoming limitations of conventional linear approaches. Results reveal that the U.S.-listed FLMB ETF acts as the strongest information source, the Europe-listed KLMH.F serves as the largest information sink, and HGGB exhibits significant cross-market propagation power across all three regions; overall, the U.S. market dominates the information network. This work provides novel empirical evidence and a methodological framework for understanding financial integration and systemic risk transmission in green bond markets.

0 citationsRead paper