AnyJev Technical Report

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses label prior bias and positional order effects in typed decision-making with pretrained language models by proposing a gradient-free correction method based on single-pass prefilling. Specifically, the approach eliminates biases through Bayesian prior correction and cyclic rotation averaging of log-probabilities, while introducing an adaptive early stopping strategy over unlabeled data to optimize inference efficiency. Requiring no parameter updates, the method integrates vLLM acceleration with dynamic sampling, significantly reducing order-flip rates and improving accuracy. Furthermore, it decreases the number of required prefill passes to approximately 7.3 and increases serving throughput by 2.2×, achieving an effective balance between precision and computational efficiency.
📝 Abstract
A typed decision is a choice among a fixed set of options, returned as a probability rather than as text. Systems that need typed decisions today use models trained for that purpose. This report describes AnyJev, which reads a typed decision from one prefill of a pretrained instruction-tuned language model. The readout restricts the next-token distribution at the answer position to the option tokens. It has two defects: the model assigns higher probability to some labels whatever the input, and to some positions in the option list. AnyJev corrects both with no gradient steps and no parameter changes: it divides out a label prior estimated from unlabelled inputs, and it averages log-probabilities over the K cyclic rotations of the option list. On two 20-option tasks the rotations lower the order-flip rate from 0.33 to 0.14 and from 0.33 to 0.18, and raise accuracy on 11 of 11 models on both. Reading every rotation requires K prefills. A stopping rule selected against the full-rotation decision on unlabelled states cuts that. Selecting the threshold on one unlabelled split and bounding its disagreement on a second, it reads 10.6 rotations of 18 at a verified 0.008 bound on two of four cells; selected and bounded on one split, as our serving run did, it reads 7.3 and serves 2.2 times as many decisions per second on vLLM. The code is open source.
Problem

Research questions and friction points this paper is trying to address.

typed decision
label bias
position bias
language model
probability readout
Innovation

Methods, ideas, or system contributions that make the work stand out.

typed decision
label prior correction
cyclic rotation averaging
training-free
early stopping rule