Score
Designs and implements adaptive score-distillation-sampling (SDS) algorithms that produce or optimize outputs consistent across multiple viewpoints by modulating per-view attention, aligning instance and spatial tokens, and applying per-view adaptive weighting or regularization to suppress cross-view hallucinations during optimization. Analyzes and evaluates multi-view consistency and failure modes of such SDS-based optimization procedures to guide attention modulation and alignment strategies.
Single-view score distillation in 3D generation suffers from high gradient variance and global shape inconsistency. This work proposes Multi-View Score Distillation Interpolation (MV-SDI), which aggregates gradients from antipodal view pairs to substantially reduce gradient estimation variance along the camera axis, while keeping the pre-trained 2D diffusion model frozen, requiring no multi-view data, and incurring no additional peak memory cost. Under a fixed UNet evaluation budget, MV-SDI with only K=2 views improves CLIP R-Precision to 83.8% and halves the number of optimization steps; with K=4 views, it reduces the step count by fourfold and achieves an R-Precision of 86.9%, significantly outperforming single-view baselines across all alignment metrics.
To address the challenge of fine-grained user-intent alignment in Score Distillation Sampling (SDS)—particularly in text-to-3D generation—this work introduces, for the first time, a learnable reward model into the SDS framework. We propose a reward-weighted noise sampling mechanism and a corresponding weighted SDS loss. Extending this to a variational setting, we develop RewardVSD, enabling consistent optimization across multiple reward dimensions. Our method unifies pretrained diffusion priors, end-to-end differentiable reward modeling, and variational score distillation, achieving gradient-level intent alignment. Experiments demonstrate that RewardVSD consistently outperforms both SDS and Variational Score Distillation (VSD) across text-to-image generation, 2D image editing, and text-to-3D synthesis. It establishes new state-of-the-art performance in generation quality, semantic fidelity, and controllability.
In text-conditioned image/3D generation, Score Distillation Sampling (SDS) often yields blurry edits and identity distortion (e.g., pose or structural misalignment) due to noisy gradient estimates. To address this, we propose an identity-preserving distillation sampling framework centered on Fixed-Point Regularization (FPR)—the first method to directly regularize the text-conditioned score function within the SDS paradigm, enabling self-calibration of gradient bias without requiring reference image pairs and thereby ensuring identity consistency before and after editing. Our approach significantly improves structural fidelity and detail sharpness in both text-driven image editing and editable Neural Radiance Fields (NeRFs), effectively suppressing blur and identity drift. Quantitative and qualitative evaluations demonstrate consistent superiority over existing state-of-the-art methods across multiple metrics.
Traditional Score Distillation Sampling (SDS) for 3D generation often suffers from texture oversaturation and geometric distortion; while negative prompting mitigates these issues, it introduces an inherent trade-off between texture enhancement and shape fidelity. This paper proposes Target-aware Multi-Objective Score distillation (T-MOS), the first framework to characterize the coupled influence of target-embedding-based negative prompts on both texture and geometry. T-MOS introduces an adaptive weighting mechanism that dynamically balances texture realism and geometric accuracy during optimization. Built upon pretrained 2D text-to-image diffusion models, it requires no auxiliary networks or explicit supervision. Extensive experiments demonstrate that T-MOS consistently outperforms state-of-the-art methods across multiple benchmarks, generating 3D assets with both high-fidelity textures and precise geometric structures.
Generative super-resolution (GSR) models improve perceptual quality but often introduce perceptually inconsistent “hallucinated” details—artifacts misaligned with either the low-resolution input or ground-truth high-resolution images—hindering real-world deployment. Method: We propose the first hallucination quantification metric based on multimodal large language models (MLLMs), yielding scores highly correlated with human subjective assessments (Pearson’s *r* > 0.92). To mitigate hallucination, we design a differentiable deep feature distance as a reinforcement learning reward signal to enforce input-output semantic consistency in the generator. Results: Our approach significantly suppresses hallucination (average reduction of 37.6%) while preserving fidelity. Crucially, the MLLM-based hallucination score is complementary to conventional metrics (e.g., LPIPS, NIQE), enabling more holistic GSR evaluation and optimization. This work establishes a new paradigm for hallucination-aware GSR assessment and training.
Large Vision-Language Models are prone to hallucination due to cross-modal attention imbalance. This work proposes FLASH, a framework that corrects visual information flow via spectral surgery to mitigate this issue. Specifically, we identify two novel hallucination patterns and design a Spectral Vortex Score to localize anomalous visual attention heads, followed by adaptive frequency-domain local attention shaping. The proposed framework effectively alleviates hallucinations without requiring additional training or contrastive decoding, thereby preserving inference efficiency. Comprehensive evaluations demonstrate that FLASH achieves superior overall performance compared to existing state-of-the-art methods.
This work addresses the susceptibility of large vision-language models (LVLMs) to hallucination, which arises from visual information degradation during perception and memory phases. To mitigate this issue without requiring additional training, the authors propose a saliency-driven perception realignment (SDPR) framework that jointly optimizes attention, key-value (KV) cache, and decoding throughout inference. Specifically, SDPR employs saliency-guided attention redistribution to recover critical visual evidence, spatially aligns the KV cache to preserve relevant features, and incorporates prior-constrained contrastive decoding to suppress language prior bias. This approach is the first to systematically alleviate LVLM hallucination from the perspective of visual perception degradation, achieving state-of-the-art performance across diverse architectures on both hallucination benchmarks and general multimodal evaluations, while incurring minimal inference overhead.
This work addresses the perceptual bottlenecks commonly observed in large vision-language models, which often manifest as attentional bias and insufficient robustness under image degradation. To mitigate these issues, the authors propose P²-DPO, a novel training paradigm built upon the Direct Preference Optimization (DPO) framework. P²-DPO introduces, for the first time, a perception-oriented online self-generated preference pair mechanism and incorporates a calibration loss to achieve causal alignment between vision and language modalities. Notably, this approach operates without human feedback and, at comparable training cost, substantially enhances model performance in terms of attention region fidelity and robustness in degraded visual conditions, thereby improving both perceptual accuracy and visual reliability.
Diffusion models often generate images with hallucinations due to oversmoothed score functions, compromising output reliability. This work formally establishes, for the first time, a theoretical link between score smoothness and the probability mass of hallucinations from a data density perspective. To mitigate this issue, we propose Variance-guided Score Modulation (VSM), a strategy that reduces score oversmoothing by explicitly regulating the Jacobian of the score function to better approximate the true score. We introduce two benchmark datasets exhibiting extreme semantic variation to enable systematic evaluation. Experimental results demonstrate that VSM consistently reduces hallucinations by approximately 25% across multiple datasets while preserving high fidelity and diversity in generated images, thereby significantly enhancing the trustworthiness of diffusion models.
This work addresses the susceptibility of large vision-language models (LVLMs) to hallucination during generation, which undermines their reliability. To mitigate this issue, the authors propose LTS-FS, a plug-and-play framework that introduces, for the first time, a layer-wise sparsity control strategy grounded in causal intervention-based attribution. By quantifying the causal contribution of each network layer to hallucinatory outputs, LTS-FS dynamically modulates the strength of feature guidance and applies precise interventions only to layers highly associated with hallucination. This targeted approach avoids perturbing irrelevant layers, thereby effectively suppressing hallucinations while preserving the model’s performance on general tasks. Extensive experiments demonstrate that LTS-FS significantly alleviates hallucination across multiple LVLMs and benchmarks without requiring any retraining.