Robust Explanations for User Trust in Enterprise NLP Systems

📅 2026-04-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of effective evaluation for the robustness of token-level explanations in real-world enterprise NLP systems deployed as black boxes under authentic user noise, which undermines user trust. The work proposes the first unified black-box robustness evaluation framework, integrating leave-one-out masking with multiple realistic perturbations—substitution, deletion, shuffling, and back-translation—and introduces the top-token flip rate as a key metric. Large-scale experiments across six models, including BERT, RoBERTa, Qwen, and Llama, enable the first systematic cross-architecture comparison of explanation robustness between encoders and decoder-based large language models (LLMs). Results reveal that decoder LLMs significantly outperform encoder models, exhibiting 73% lower average flip rates, and demonstrate that scaling model size from 7B to 70B parameters improves stability by 44%, alongside a derived cost–robustness trade-off curve.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Adversarial Learning & RobustnessComputer Vision: Adversarial Attacks & Robustness

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common case of black-box deployment (API-only access) where representation-based explainers are infeasible and existing studies provide limited guidance on whether explanations remain stable under real user noise, especially when organizations migrate from encoder classifiers to decoder LLMs. To close this gap, we propose a unified black-box robustness evaluation framework for token-level explanations based on leave-one-out occlusion, and operationalize explanation robustness with top-token flip rate under realistic perturbations (swap, deletion, shuffling, and back-translation) at multiple severity levels. Using this protocol, we conduct a systematic cross-architecture comparison across three benchmark datasets and six models spanning encoder and decoder families (BERT, RoBERTa, Qwen 7B/14B, Llama 8B/70B; 64,800 cases). We find that decoder LLMs produce substantially more stable explanations than encoder baselines (73% lower flip rates on average), and that stability improves with model scale (44% gain from 7B to 70B). Finally, we relate robustness improvements to inference cost, yielding a practical cost-robustness tradeoff curve that supports model and explanation selection prior to deployment in compliance-sensitive applications.
Problem

Research questions and friction points this paper is trying to address.

robust explanations
user trust
enterprise NLP
black-box deployment
explanation stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

black-box robustness
token-level explanations
explanation stability
large language models
cost-robustness tradeoff
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Guilin Zhang
Workday AI
K
Kai Zhao
Workday AI
J
Jeffrey Friedman
Workday AI
X
Xu Chu
Workday AI
A
Amine Anoun
Workday AI
J
Jerry Ting
Workday AI