🤖 AI Summary
This study addresses label prior bias and positional order effects in typed decision-making with pretrained language models by proposing a gradient-free correction method based on single-pass prefilling. Specifically, the approach eliminates biases through Bayesian prior correction and cyclic rotation averaging of log-probabilities, while introducing an adaptive early stopping strategy over unlabeled data to optimize inference efficiency. Requiring no parameter updates, the method integrates vLLM acceleration with dynamic sampling, significantly reducing order-flip rates and improving accuracy. Furthermore, it decreases the number of required prefill passes to approximately 7.3 and increases serving throughput by 2.2×, achieving an effective balance between precision and computational efficiency.
📝 Abstract
A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout restricts the next-token distribution at the answer position to the option tokens. It has two defects: the model assigns higher probability to some labels whatever the input, and to some positions in the option list. AnyJev corrects both with no gradient steps and no parameter changes: it divides out a label prior estimated from unlabelled inputs, and it averages log-probabilities over the K cyclic rotations of the option list. On two 20-option tasks the rotations lower the order-flip rate from 0.33 to 0.14 and from 0.33 to 0.18, and raise accuracy on 11 of 11 models on both. Reading every rotation requires K prefills. A stopping rule selected against the full-rotation decision on unlabelled states cuts that. Selecting the threshold on one unlabelled split and bounding its disagreement on a second, it reads 10.6 rotations of 18 at a verified 0.008 bound on two of four cells; selected and bounded on one split, as our serving run did, it reads 7.3 and serves 2.2 times as many decisions per second on vLLM. The code is open source.