Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability of text-to-3D evaluation leaderboards caused by rendering configuration discrepancies, which induce score fluctuations that obscure genuine model performance differences. By fixing scenes while systematically varying rendering and prompting factors, the authors quantify evaluator sensitivity to configurations. They introduce the concept of a "protocol margin envelope" to decouple scoring stability, decision uncertainty, and sensitivity diagnostics, supported by statistical inference through large-scale ablation studies and multi-evaluator benchmarking. The findings reveal that while winning rankings for most evaluators shift across configurations, no conclusive rank reversals are observed. Consequently, this work recommends that future research explicitly report protocol-dependent comparative results, providing critical guidance for constructing more robust evaluation frameworks in 3D generation.
📝 Abstract
Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors for 19 alignment evaluators plus one perceptual-quality control, then test four targeted scene degradations. Peak configuration variance exceeds between-generator variance for 17/19 alignment evaluators, with prompt-bootstrap lower bounds above 1 for 11/19. Rankings are more stable than scores, yet 18/19 evaluators change their point-estimate winner under some configuration. Pairwise protocol margin envelopes show which comparisons keep their direction across the tested settings. Selected pairs have opposite pointwise intervals, but no reversal survives simultaneous inference over the full search. Thus the observed winner changes are descriptive, not confirmed changes in generator superiority. Sensitivity remains separate: no evaluator, even the prompt-free control, exceeds 67% tie-adjusted directional discrimination on layout scrambling, which is diagnostic rather than human-validated ground truth. The audit separates score stability, decision uncertainty, and targeted sensitivity, and recommends reporting (generator, score, card ID) with protocol-dependent comparisons and selection-aware uncertainty.
Problem

Research questions and friction points this paper is trying to address.

text-to-3D evaluation
configuration fragility
leaderboard stability
rendered-image evaluation
alignment evaluators
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-to-3D Evaluation
Configuration Fragility
Rendered-image Audit
Decision Uncertainty
Sensitivity Analysis
🔎 Similar Papers