π€ AI Summary
This study addresses the limited interaction capacity of linear adapters in frozen language models and the parameter redundancy inherent in explicit second-order methods by proposing the SQUARE adapter. This approach leverages variational quantum circuits to perform amplitude encoding and Pauli-Z measurements on bottleneck vectors, theoretically demonstrating that its outputs strictly correspond to normalized quadratic forms. Consequently, it efficiently models O(dΒ²) feature interactions using a minimal set of shared parameters. Combined with PyTorch-based batched simulation techniques, the method enables training without requiring actual quantum hardware. Evaluated on GLUE benchmarks, SQUARE achieves an average accuracy of 0.7565, significantly outperforming affine quadratic predictors and MLP baselines while maintaining its superiority in few-shot scenarios.
π Abstract
Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain linear at the adaptation module itself, whereas explicit second-order alternatives introduce pairwise interactions through direct parameterization or predefined factorizations. We propose SQUARE, a Structured QUAntum REpresentation adapter that amplitude-encodes the bottleneck vector, applies a parameterized quantum circuit, and measures the resulting state. We show that each basis-probability feature is exactly a normalized quadratic form in the bottleneck coordinates, while the additional Pauli-$Z$ readouts are signed linear combinations of these probabilities. The measured map can therefore parameterize interactions over $O(d^2)$ coordinate pairs through a small set of shared circuit parameters, where $d$ is the bottleneck dimension. It provides a structured parameterization within, rather than beyond, the classical normalized-quadratic feature class. In a disjoint same-pipeline evaluation over eight GLUE-derived controlled interaction tasks and five shared seeds, SQUARE achieves an average test accuracy of $0.7565$, compared with $0.7355$ for an affine normalized-quadratic predictor, $0.7271$ for the evaluated parameter-matched Givens mixing model, $0.6817$ for an MLP, and $0.6155$ for a frozen-circuit control. Under reduced supervision, it also shows consistent gains over the strongest evaluated classical comparator, with the same qualitative pattern across multiple frozen LM backbones. All circuit experiments use simulation, while the learned feature map can be evaluated exactly in batched PyTorch without quantum hardware.