Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability
论文提出了一种反向推理模式构建方法,通过两阶段数学数据集训练和细粒度奖励机制,提高大型语言模型的反向思维能力和动态适应性,以解决复杂问题。
论文提出了一种反向推理模式构建方法,通过两阶段数学数据集训练和细粒度奖励机制,提高大型语言模型的反向思维能力和动态适应性,以解决复杂问题。
本文针对稀疏视角下3D高斯点渲染产生的几何不一致问题,提出了一种利用预训练图像扩散模型生成伪视图并进行选择性优化的方法,以提高重建质量和可靠性。
Existing sparse-view 3D editing methods rely on test-time iterative optimization, resulting in high computational costs, cross-view inconsistency, and limited generalization. This work proposes a feed-forward 3D editing framework that eliminates the need for per-scene optimization at test time by incorporating cross-view image-domain regularization and geometric alignment constraints during training. Leveraging text-guided editing, multi-view joint supervision, and a 3D Gaussian splatting representation, the method generates consistent and high-fidelity 3D content without scene-specific refinement. The approach significantly improves cross-view consistency, achieves inference speeds several orders of magnitude faster than existing methods, and maintains high editing fidelity.
This work proposes an end-to-end feedforward anchored scene Transformer architecture to address the limitations of existing 3D instance segmentation methods, which predominantly rely on non-end-to-end “lift-and-cluster” paradigms that decouple representation learning from segmentation objectives and hinder scalability. The proposed approach introduces learnable 3D anchor generation and anchor-sampling cross-attention mechanisms to achieve multi-view consistent instance segmentation without post-hoc clustering. To mitigate query conflicts and enhance boundary precision, it incorporates dual-level regularization, multi-view contrastive learning, and a dynamic spatial overlap penalty. Evaluated on complex indoor scene datasets, the method significantly outperforms current clustering-based baselines in segmentation accuracy, memory efficiency, and inference speed.
In visual localization, Absolute Pose Regression (APR) models often suffer from limited generalization due to end-to-end black-box learning that lacks explicit 3D geometric understanding. To address this, we propose Geometric Representation Regression (GRR), a novel paradigm that abandons direct regression of 6-DoF poses. Instead, GRR separately regresses ray direction bundles (encoding rotation) and point maps (encoding translation), and integrates a differentiable deterministic geometric solver for end-to-end joint optimization in the world coordinate system. Crucially, GRR is the first method to leverage the inverse process of novel view synthesis for pose estimation—explicitly decoupling geometric representation learning from pose solving while embedding strong geometric priors. Evaluated on 7-Scenes and Cambridge Landmarks, GRR achieves state-of-the-art performance, delivering significant improvements in both absolute pose accuracy and cross-scene robustness.
论文提出了一种反向推理模式构建方法,通过两阶段数学数据集训练和细粒度奖励机制,提高大型语言模型的反向思维能力和动态适应性,以解决复杂问题。
本文针对稀疏视角下3D高斯点渲染产生的几何不一致问题,提出了一种利用预训练图像扩散模型生成伪视图并进行选择性优化的方法,以提高重建质量和可靠性。
Existing sparse-view 3D editing methods rely on test-time iterative optimization, resulting in high computational costs, cross-view inconsistency, and limited generalization. This work proposes a feed-forward 3D editing framework that eliminates the need for per-scene optimization at test time by incorporating cross-view image-domain regularization and geometric alignment constraints during training. Leveraging text-guided editing, multi-view joint supervision, and a 3D Gaussian splatting representation, the method generates consistent and high-fidelity 3D content without scene-specific refinement. The approach significantly improves cross-view consistency, achieves inference speeds several orders of magnitude faster than existing methods, and maintains high editing fidelity.
This work proposes an end-to-end feedforward anchored scene Transformer architecture to address the limitations of existing 3D instance segmentation methods, which predominantly rely on non-end-to-end “lift-and-cluster” paradigms that decouple representation learning from segmentation objectives and hinder scalability. The proposed approach introduces learnable 3D anchor generation and anchor-sampling cross-attention mechanisms to achieve multi-view consistent instance segmentation without post-hoc clustering. To mitigate query conflicts and enhance boundary precision, it incorporates dual-level regularization, multi-view contrastive learning, and a dynamic spatial overlap penalty. Evaluated on complex indoor scene datasets, the method significantly outperforms current clustering-based baselines in segmentation accuracy, memory efficiency, and inference speed.
In visual localization, Absolute Pose Regression (APR) models often suffer from limited generalization due to end-to-end black-box learning that lacks explicit 3D geometric understanding. To address this, we propose Geometric Representation Regression (GRR), a novel paradigm that abandons direct regression of 6-DoF poses. Instead, GRR separately regresses ray direction bundles (encoding rotation) and point maps (encoding translation), and integrates a differentiable deterministic geometric solver for end-to-end joint optimization in the world coordinate system. Crucially, GRR is the first method to leverage the inverse process of novel view synthesis for pose estimation—explicitly decoupling geometric representation learning from pose solving while embedding strong geometric priors. Evaluated on 7-Scenes and Cambridge Landmarks, GRR achieves state-of-the-art performance, delivering significant improvements in both absolute pose accuracy and cross-scene robustness.