FORGE: Verification-Gated Behavioral Repair for Generative Language Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of mitigating post-deployment bias and toxicity in large language models (LLMs) while providing certified repairs that preserve original capabilities. To this end, it proposes FORGE, a novel decoupled architecture that disentangles defect localization, weight editing, and behavioral verification. By integrating constrained quadratic optimization, null-space projection, and causal probing techniques with a repair abstraction mechanism for edit-agnosticism, the framework achieves per-sample correctness guarantees for autoregressive generation. Extensive evaluations across five open-source LLMs demonstrate that FORGE substantially reduces bias and toxicity with minimal perplexity degradation, outperforming conventional fine-tuning approaches particularly in few-shot scenarios.
📝 Abstract
Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally, provide correctness guarantees. Existing approaches address this only partially: gradient-based fine-tuning lacks per-instance guarantees and becomes unstable with few defect samples; model editing assumes explicit knowledge replacement rather than behavioral correction; and constraint-based repair is largely restricted to discriminative models with unique target outputs. We present FORGE, a framework for targeted behavioral repair of generative language models that separates defect localization, weight editing, and behavioral verification into independent stages. Its core is a repair abstraction that converts localized defective generation into explicit optimization objectives, enabling verification-oriented repair techniques to operate on autoregressive generation. FORGE is editing-mechanism agnostic: we instantiate it with (1) a constraint-based quadratic optimization method that provides per-sample repair certificates and (2) a null-space projection editor that minimizes interference with the original model distribution, both under the same localization and verification protocol. On five open-source LLMs, FORGE consistently achieves larger reductions in bias and toxicity than gradient-based fine-tuning with minor perplexity degradation. The two backends exhibit complementary performance across architectures, which a lightweight causal probe traces to where toxicity-related signals concentrate. FORGE also remains effective with only a handful of defective examples, where conventional fine-tuning often oscillates or fails to converge.
Problem

Research questions and friction points this paper is trying to address.

generative large language models
behavioral repair
demographic bias
toxic generation
model editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

behavioral repair
generative language models
verification-gated
constraint-based optimization
null-space projection
Hsin-Ling Hsu
Hsin-Ling Hsu
National Chengchi University
Information RetrievalNatural Language ProcessingAI for HealthcareTrustworthy AI
M
Min-Yu Chen
Department of Management Information Systems, National Chengchi University, Taipei, Taiwan
N
Nai-Chia Chen
Department of Management Information Systems, National Chengchi University, Taipei, Taiwan
Y
Yan-Ru Chen
Department of Management Information Systems, National Chengchi University, Taipei, Taiwan
Y
Yi-Ling Chang
Department of Management Information Systems, National Chengchi University, Taipei, Taiwan
Fang Yu
Fang Yu
Associate Professor, Dept. Management Information Systems, National Chengchi University
Software VerificationString AnalysisAutomata TheoryWeb Security