Score
Designs and executes controlled input perturbations and targeted adversarial modifications to probe model behavior and surface defects. Builds guided fuzzers and experiment protocols, analyzes model outputs to iteratively refine perturbation strategies and hypotheses, and documents and catalogs confirmed defects.
Machine learning training is significantly more time-consuming than inference, and the design of input or parameter perturbations has long relied on empirical trial-and-error. Method: This paper models training dynamics as a first-passage process and introduces a statistical mechanics framework to analyze model responses to input/parameter perturbations. It proposes, for the first time, a single-frequency perturbation response theory grounded in the quasi-stationary assumption, and rigorously proves its generalizability to multi-frequency perturbation regimes—enabling rational optimization of perturbation protocols. Contribution/Results: Evaluated on ResNet-18 trained for CIFAR-10 classification, the method precisely identifies the optimal perturbation type and frequency, reducing training iterations by 23% and improving test accuracy by 1.4 percentage points, thereby substantially enhancing both training efficiency and generalization performance.
Defects in deep learning frameworks pose severe security risks in safety-critical domains; however, existing fuzzing techniques underutilize multi-source feedback and suffer from coarse granularity and low automation. This paper proposes FUEL—the first feedback-driven fuzzing framework leveraging dual large language model (LLM) agents: an *analysis LLM* performs fine-grained interpretation of coverage, crashes, and anomalies, while a *generation LLM* evolves high-diversity test cases based on this feedback, enabling closed-loop, synergistic feedback utilization. FUEL overcomes the static and unidirectional nature of conventional fuzzing feedback mechanisms. Evaluated on PyTorch and TensorFlow, FUEL identified 104 vulnerabilities, including 93 previously unknown ones; 47 have been patched, and 5 have received CVE identifiers.
This study addresses the tendency of large language models to circumvent alignment objectives through superficial compliance, resulting in internal representations that fail to genuinely internalize safe behaviors. To overcome this limitation, we propose a probe-guided fine-tuning approach that, for the first time, employs continuously updated internal probes as direct optimization signals. By leveraging both linear and nonlinear probing techniques, our method shapes internal representations specifically for harmlessness and honesty, transcending the constraints of relying solely on output-level feedback. Empirically, this approach significantly outperforms Direct Preference Optimization (DPO) and inference-time interventions in navigating the safety-utility trade-off. It substantially enhances robustness against jailbreak attacks while preserving the linear encoding of concepts to ensure continued monitorability.
This work addresses the inefficiency of existing deep neural network fuzzing methods, which struggle to effectively explore high-dimensional, heterogeneous input spaces due to per-sample iteration and uniform perturbation strategies. The authors propose a tensor-based batch fuzzing framework that embeds input constraints and output property checks as non-trainable network layers, enabling specification-aware parallel testing. By integrating adaptive perturbation scaling—either isotropic or anisotropic—driven by norm-bounded feasible regions, the method processes multiple constrained inputs simultaneously within a single batched iteration. This approach significantly enhances both exploration efficiency and precision. Experimental results demonstrate up to a 40× throughput improvement and a 4× increase in violation detection across three major benchmarks, substantially outperforming conventional sequential fuzzing techniques.
In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
为解决JavaScript引擎中的高影响漏洞问题,提出StateLens框架,利用大型语言模型智能选择高价值的插桩目标,指导模糊测试探索未被发现的引擎语义。
研究使用线性探针检测大型语言模型中工具调用错误的有效性,通过18个模型测试,发现该方法能有效识别多种错误。
This study addresses the proliferation of invalid reports in deep learning compiler fuzzing caused by repeatedly triggering known bugs. To this end, we propose Reprise, a tool that departs from traditional post-hoc deduplication paradigms by distilling known bugs into semantic graph patterns and proactively avoiding their trigger points during test case synthesis for source-level deduplication. Leveraging a unified intermediate representation and local operator signatures, Reprise employs a lightweight generator guided by semantic graph pattern matching to steer graph synthesis, thereby suppressing redundant bug triggers prior to compilation. Experimental results demonstrate that Reprise reduces crash reports by 90.2% and 93.8% for TVM and Inductor, respectively, while maintaining code coverage and uncovering 25 previously unknown bugs.
该研究通过利用大型语言模型自动生成注释来指导模糊测试,以解决依赖人工专家注释的可扩展性问题,并评估了这种方法在漏洞检测中的有效性。