Score
Designs and implements methods that detect unstable or low-confidence tokens in a model’s generated sequence by applying perturbations, repeated inference, or other stability tests, and quantify token-level uncertainty. Builds inference-time rewriting procedures that selectively propose, score, and apply token revisions (perturbation-based candidate generation and replacement) to repair or improve the final decoded output without full re-generation.
This work addresses the challenge in block-wise parallel speculative decoding where localized high uncertainty in long-sequence generation leads to clustered errors, resulting in elevated rejection rates and diminished speedup. To mitigate this, the authors propose a budget-aware dynamic repair tree structure that identifies uncertainty hotspots at prediction confidence boundaries and constructs bounded repair paths only at critical nodes. A lightweight state resynchronization mechanism is further introduced to enhance acceptance length and verification efficiency with minimal overhead. The method seamlessly integrates with mainstream parallel draft generation frameworks and demonstrates consistent improvements across multiple code and mathematical reasoning benchmarks, achieving a 4.2–7.5% increase in average acceptance length and end-to-end inference speeds of 2.66–3.49× over baseline autoregressive decoding.
This work addresses the challenge of dynamic reasoning failures in large language models—such as “mid-reasoning distraction”—which often lead to incorrect outputs but remain undetected by conventional evaluations that focus solely on final answers. To capture such process-level breakdowns, the authors propose a training-free, model-agnostic diagnostic method that leverages token-level log probabilities available during inference. By combining Jensen-Shannon divergence and entropy, they construct a dynamic instability metric that distinguishes between “constructive” and “destructive” instability, revealing how the timing of instability critically influences the model’s capacity for self-correction. Experiments on GSM8K and HotpotQA demonstrate that peak instability intensity strongly predicts reasoning errors (with AUC significantly above random chance) and exhibits a monotonic decline in accuracy as model scale increases, confirming the method’s effectiveness and scalability.
Quantifying output variability of large language models (LLMs) under input perturbations or model variants remains challenging due to the intractability of explicit probability distribution modeling and sensitivity to stochasticity. Method: This paper proposes a black-box robustness auditing framework that formulates output divergence detection as a statistical hypothesis test in semantic embedding space (e.g., BERTScore, STS). It constructs empirical null distributions via Monte Carlo sampling—bypassing explicit modeling of output distributions and mitigating randomness-induced bias. Contribution/Results: We introduce the first distributed perturbation analysis paradigm, enabling model-agnostic, multi-perturbation joint testing, interpretable p-values, scalar effect sizes, and integrated multiple-testing correction (e.g., Bonferroni). Experiments demonstrate accurate quantification of response shifts, reliable estimation of true/false positive rates, and cross-model consistency assessment—achieving significantly improved reliability and reproducibility in LLM robustness evaluation without distributional assumptions.
为提高大型语言模型在下游任务中的推理能力,提出基于稳定性的测试时适应方法TASCO,通过优化局部稳定性来改善模型的自信度和准确性。
Autoregressive language models suffer from error accumulation due to their unidirectional generation mechanism. To address this, we propose Resample-Previous-Tokens (RPT), the first plug-and-play local resampling method integrated into standard autoregressive decoding—without modifying the model architecture. RPT iteratively backtracks and resamples previously generated tokens within a sliding window, enabling inference-time correction in a zero-fine-tuning setting; it also supports lightweight fine-tuning (using only ~100B tokens) for further gains. Evaluated on an 8B-parameter model, RPT achieves approximately 10% relative improvement on both programming and general reasoning benchmarks. It effectively mitigates error propagation while preserving decoding efficiency, striking a favorable balance between correction capability and computational overhead.
This study addresses the overconfidence problem in large language models (LLMs) for code generation, where erroneous programs frequently elicit confidence levels comparable to correct ones. Leveraging open-source models and execution-based benchmarks, we systematically evaluate uncertainty metrics—including global and local token-level confidence, entropy, and latent space representations—and assess the effectiveness of conventional mitigation strategies. Our findings reveal that instruction tuning exacerbates spurious certainty in LLMs. Furthermore, existing uncertainty metrics only partially capture execution failures, and standard mitigation approaches prove unreliable in addressing this issue. Notably, however, our analysis suggests that latent representations may encode correctness signals not explicitly surfaced during decoding, highlighting a promising direction for future research toward more robust uncertainty estimation in neural code generation.
研究通过分析输出行为、隐藏状态几何和注意力头功能,探讨六种自然和合成输入扰动在大型语言模型中的传播情况,揭示单一指标评估鲁棒性的局限性。
This work identifies two distinct failure modes in large language model reasoning: decisive failures and persistently uncertain failures. It proposes the first fine-grained diagnostic method based on token-level uncertainty signals to detect verifiable signatures of these failure types within reasoning trajectories, thereby revealing the dynamic boundary of failure detectability. By constructing a cross-model and cross-dataset validation framework, the study reproduces these failure signatures across 23 configurations, with 20 showing statistically significant alignment with predictions—far exceeding random chance. Leveraging these insights, the authors adaptively refine self-consistency strategies, substantially improving failure detection performance.
This work addresses the vulnerability of large language models to short-sequence injection attacks at intermediate positions during text generation—a risk inadequately mitigated by existing alignment methods that focus solely on initial outputs. To ensure safety throughout the entire generation process, the authors propose a trajectory-based alignment mechanism that simulates intermediate perturbations by injecting adversarial sequences during reasoning. This approach integrates rejection-direction alignment with adversarial training to enforce consistent safety across all generation steps. Experimental results demonstrate that the method substantially enhances robustness against mid-generation injection attacks and generalizes effectively to early-stage attack scenarios, thereby validating that process-level alignment outperforms conventional paradigms relying only on final outputs or hidden states.
This work addresses the sensitivity of large language models to minor input perturbations, a challenge inadequately handled by existing methods that enforce sequence-level consistency and fail to capture localized semantic drifts. To this end, the authors propose the S²R² framework, which introduces a segment-level robustness mechanism during LoRA fine-tuning. S²R² decomposes outputs into semantic segments and aligns clean and perturbed generations via optimal transport, selectively penalizing segments exhibiting maximal drift. Additionally, it incorporates an adapter stability regularizer motivated by attention reallocation, constraining LoRA norms to mitigate evidence shift. Theoretical analysis from a PAC-Bayesian perspective reveals that controlling adapter growth enhances generalization across perturbations. Experiments demonstrate that S²R² significantly improves robustness against spelling errors, deletions, synonym substitutions, and paraphrasing in summarization tasks, while preserving strong performance on clean inputs and outperforming consistency-based baselines in cross-dataset transfer.