π€ AI Summary
This study addresses the disconnect between high image fidelity and screen content readability in 3D Gaussian Splatting (3DGS) by constructing the first static screen content benchmark. Leveraging procedural scene generation, multi-view segmentation, and optical character recognition (OCR), it enables decoupled evaluation of overall fidelity, screen region quality, and text readability. Experiments reveal significant inconsistencies between screen PSNR and OCR-based rankings across most methods, exposing the risk of relying solely on PSNR for model selection. By providing a controllable testing protocol with precise annotations, this work bridges the gap in existing 3DGS evaluation frameworks regarding functional readability metrics.
π Abstract
Does high image fidelity imply readable screen content in 3D Gaussian Splatting (3DGS)? We introduce 3DGS-SC, a controlled static screen-content dataset and benchmark for examining this mismatch. Ten procedural scenes provide fixed multi-view splits, exact cameras, screen masks, text boxes, and transcripts. The protocol separates whole-image fidelity, screen-region fidelity, OCR readability, and edge preservation. In the reported five-method comparison, LightGaussian exceeds Mip-Splatting in screen PSNR by only 0.08 dB, yet trails it in OCR accuracy by 12.7 percentage points. Across all ten method pairs, screen-PSNR and OCR orderings disagree in seven cases. These aggregate results expose a method-selection failure of fidelity-only evaluation. Complementing prior text-aware 3DGS research, 3DGS-SC targets controlled monitor interfaces and exact annotations; scene-wise robustness and acquisition effects remain open validation questions.