StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing approaches to indoor furniture arrangement often suffer from style inconsistencies in shape, material, and color due to either independent selection or reliance on static, local relationships. To address this, this work proposes StyleForge, a framework that models high-order furniture dependencies through a dynamic hypergraph-based style field. It leverages a frozen multimodal large language model to extract stylistic priors and integrates counterfactual preference learning with a Mahalanobis distance-based energy function to enable context-aware, scene-level style harmonization. The method supports test-time adaptation by updating only room-specific candidate distributions to resolve cross-slot conflicts. Evaluated on the 3D-FRONT dataset, StyleForge significantly outperforms existing object-level and scene-level retrieval baselines, achieving state-of-the-art performance in both furniture retrieval accuracy and overall stylistic coherence.
📝 Abstract
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.
Problem

Research questions and friction points this paper is trying to address.

indoor furniture styling
fixed-layout
style coherence
scene composition
furniture arrangement
Innovation

Methods, ideas, or system contributions that make the work stand out.

counterfactual reasoning
hypergraph field
style coherence
structured selection
test-time training
🔎 Similar Papers
No similar papers found.