Score
Designs and implements conditioning mechanisms that align auxiliary features to the spatial coordinates of images or feature maps so that conditioning information corresponds directly to locations in the output. Builds spatial conditioning or aligned-feature conditioning pipelines used at training and inference to bridge domain gaps, preserve identity and appearance consistency across frames, and stabilize generated or predicted appearance.
为提高VLA模型的性能,通过引入基于距离的优势学习方法DistAL,使用嵌入空间距离作为奖励,生成更优质的价值函数,从而提升下游任务成功率。
To address insufficient generative image diversity—which hinders downstream classification model training—this paper introduces the “Augmented Conditioning” paradigm. Without fine-tuning pre-trained text-to-image diffusion models, it jointly conditions generation on both standard-augmented real images (e.g., cropping, color jittering) and textual prompts. This approach significantly enhances visual diversity of generated samples while preserving semantic fidelity and domain consistency. As a lightweight, inference-time, zero-shot, parameter-free strategy, it achieves state-of-the-art performance across five long-tailed and extreme few-shot (1-shot/5-shot) classification benchmarks. It improves average accuracy by 3.2% on long-tailed tasks and up to 18.7% in few-shot settings, with particularly notable gains in generalization to tail and rare classes.
This work addresses the limitations of existing 3D generation methods, which predominantly rely on a single modality—either images or text—and consequently struggle to simultaneously capture fine visual details and rich semantic meaning, thereby hindering precise expression of user intent. To overcome this, the paper formally introduces the task of 3D generation conditioned jointly on text and image inputs and proposes a lightweight dual-branch baseline model. The architecture employs separate backbone networks to extract features from each modality and incorporates an efficient cross-modal fusion mechanism to enable joint reasoning. Experimental results demonstrate that this bimodal approach significantly outperforms unimodal baselines in terms of generation quality, geometric fidelity, and semantic consistency, effectively validating the practical utility and complementary nature of cross-modal conditioning in 3D content creation.
For black-box models applied to complex tasks such as image segmentation, defining meaningful conditional events is challenging, leading to uncertainty estimates that fail to reflect inherent sample difficulty. Method: This paper proposes an input-dependent statistical risk control framework grounded in conformal prediction. It introduces a novel, algorithm-driven mechanism for dynamically selecting conditional function classes—bypassing manual discretization—by adaptively constructing these classes based on test-sample difficulty and integrating online parameter tuning for fine-grained, approximately conditional risk control. Contribution/Results: Experiments on regression and image segmentation demonstrate substantial improvements in uncertainty calibration accuracy. The method guarantees strict statistical risk control while enhancing generalization robustness and predictive reliability.
This work addresses the limitations of existing feedforward view synthesis methods, which rely on Plücker ray representations that are highly sensitive to camera coordinate systems, resulting in poor cross-view geometric consistency. To overcome this, the authors propose a projection-conditioning strategy that replaces raw ray inputs with 2D projection cues from the target view, effectively reformulating the task as a stable image-to-image translation problem. A tailored masked autoencoder pretraining mechanism is introduced to leverage large-scale uncalibrated data under this new conditioning paradigm. The proposed approach significantly enhances model robustness and view consistency, achieving state-of-the-art performance across multiple novel view synthesis benchmarks. Notably, it outperforms ray-based baselines by a clear margin on geometric consistency metrics, demonstrating the effectiveness of decoupling geometry representation from explicit ray parameterization.
This study addresses the degradation of high-fidelity details in image-to-3D generation caused by global encoding compression. To overcome this limitation, we propose BTC3D, a training-free inference framework that first reveals the additivity of diffusion model features. By designing hybrid tile embeddings coupled with a dynamic conditioning scheduling mechanism, our method precisely extracts and fuses local high-frequency signals, enhancing detail preservation without requiring retraining. Experimental results demonstrate that BTC3D significantly improves texture quality and visual fidelity while maintaining global structural consistency. Furthermore, it can be seamlessly integrated into existing generative pipelines.
研究通过不同几何条件训练0.8B混合语言模型,以解决物体操纵任务中的物理状态输入影响问题,但未发现训练时几何对齐的可靠优势。
本文提出GeoComposer,通过几何感知表示学习机制和强化学习策略,解决摄影构图中忽视3D场景几何一致性的问题。
本文针对3D重建模型在特定条件下的失效问题,提出了GeoCond方法,通过预测几何结构的不确定性及精修门控来提高模型的可靠性。
This study addresses the semantic mismatch and unreliable training acceleration caused by representation alignment in pixel-space diffusion models, proposing a novel method termed JAx. Establishing the principle that auxiliary supervision should optimize prediction targets rather than independent representations, JAx couples noise observations via a Markov degradation process to directly align clean image predictions across varying noise levels, replacing conventional representation alignment. By integrating an EMA teacher model, a reliability gating mechanism, and gradient variance diagnostics, it achieves prediction-space supervision without external encoders. Experiments on ImageNet demonstrate that JAx significantly reduces FID scores, accelerates convergence, and effectively suppresses small-batch gradient variance, thereby validating the superiority of prediction-space supervision.