The Hidden States Cookbook: A Large-Scale Ablation Study for Noise-Robust Conversational Intent Classification in Industry

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of intent classification accuracy and computational inefficiency caused by noisy user queries in industrial-scale dialogue systems. Leveraging the Llama-3.2 model and the BANKING77 and CLINC150 datasets, we conduct large-scale controlled ablation experiments to systematically evaluate the noise robustness of various pooling strategies and frequency-domain filtering mechanisms. Our findings demonstrate that attention pooling achieves superior performance under noisy conditions, improving the F1 score by approximately 2.8 percentage points, whereas mean pooling incurs significant degradation. Furthermore, frequency-domain filtering fails to yield consistent improvements. This work provides clear architectural selection criteria and practical guidance for developing robust intent classifiers in real-world applications.
📝 Abstract
Conversational database interfaces face a critical challenge: users naturally embed queries in conversational noise (greetings, politeness, off-topic remarks), which degrades intent classification accuracy and wastes computational resources. Despite advances in orchestration and retrieval strategies, a fundamental question remains unanswered: which pooling strategy maximizes intent classification accuracy under realistic conversational noise in production language models? This work addresses this gap through 360 controlled experiments spanning four pooling configurations (mean, max, last-token, attention, and FFT-augmented variants) using Llama-3.2-1B-Instruct on BANKING77 and CLINC150 datasets under clean/noisy conditions with ten random seeds. Key findings reveal that attention pooling consistently outperforms alternative strategies under noisy conditions (~+2.6-2.8 F1 over the default), while mean pooling degrades performance by up to ~5 F1 points. Frequency-domain filtering does not produce consistent accuracy improvements and functions primarily as a structural variation rather than an accuracy-enhancing component. These results provide concrete, evidence-based guidance for building noise-robust conversational classifiers: attention pooling is recommended for noisy interfaces, mean pooling should be avoided, and last-token pooling is appropriate for clean-query scenarios.
Problem

Research questions and friction points this paper is trying to address.

intent classification
conversational noise
pooling strategy
hidden states
noise-robust
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intent Classification
Pooling Strategy
Noise Robustness
Ablation Study
Hidden States
🔎 Similar Papers
No similar papers found.
B
Bogdan Bogachov
Department of Mechanical Engineering, McGill University, Montréal, Québec, Canada
N
Nikita Letov
Axya Inc., Montréal, Québec, Canada
Yaoyao Fiona Zhao
Yaoyao Fiona Zhao
McGill University
design and manufacturingeco-designadditive manufacturingsustainable manufacturingmachine learning application