🤖 AI Summary
This paper addresses the lack of a unified, reproducible framework for adaptive inference in Transformer models. To this end, we introduce AdaptBench—the first end-to-end open-source benchmark for adaptive inference—integrating three core techniques: progressive token pruning, sparse attention, and dynamic early exiting, enabling input-adaptive computation. The framework fully automates the GLUE evaluation pipeline, including data preprocessing, low-overhead timing, CSV-based logging, ablation studies, and joint accuracy–latency assessment. All components are modular, well-documented, and script-ready, significantly enhancing reproducibility and cross-method comparability. On SST-2, AdaptBench achieves marginally higher accuracy than optimized DistilBERT at substantially lower latency, demonstrating the efficacy and practicality of dynamic computation in low-latency NLP applications. By providing a standardized, extensible evaluation infrastructure, AdaptBench establishes a foundational benchmark for future research on adaptive Transformers.
📝 Abstract
The Efficient Adaptive Transformer (EAT) framework unifies three adaptive efficiency techniques - progressive token pruning, sparse attention, and dynamic early exiting - into a single, reproducible architecture for input-adaptive inference. EAT provides an open-source benchmarking pipeline that automates data processing, timing, and ablation across GLUE tasks (SST-2, QQP, MNLI). Although this empirical study finds that combining these mechanisms can increase latency in shallow six-layer models, it demonstrates that EAT achieves slightly higher accuracy than the optimized DistilBERT baseline on SST-2, illustrating the potential of dynamic computation for latency-sensitive NLP. The main contribution is the open, end-to-end reproducible framework - complete with scripts, CSV logging, and analysis utilities - intended to serve as a community tool for further research on adaptive transformers.