🤖 AI Summary
This work addresses the challenge of synthesizing gigapixel images from ordinary photographs and sparse microscopic close-ups at extreme magnifications up to 350×, where preserving both fine material boundary details and large-scale structural consistency is difficult. The authors propose a two-stage cascaded generative framework: the first stage recovers global pattern coherence, while the second refines local textures, guided by segmentation masks to enable reference-driven synthesis in ambiguous regions. This approach achieves, for the first time, extreme-scale super-resolution tailored to everyday objects, generating gigapixel images that exhibit both microscopic material realism and macroscopic structural consistency on a newly curated dataset. Notably, it effectively maintains global coherence of repetitive geometric patterns—such as those in fabrics—and enables fully scalable, explorable visualization across all resolutions.
📝 Abstract
We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of consumer-grade microscope close-ups, MicroZoom synthesizes a seamless, gigapixel-resolution image grounded in the material character of the real references, enabling exploratory visualization of microscopic texture across the full spatial extent of an object. Our goal is plausible synthesis, not exact reconstruction. We focus on full-image, reference-based, extreme-scale super-resolution at magnification levels of up to 350x, a setting that introduces two major challenges: (1) recovering texture-specific detail from highly lossy inputs near ambiguous material boundaries, and (2) preserving correct large-scale pattern structure, such as the repeating geometry of a fabric weave, across millions of local predictions. We address these with a two-stage cascaded design, where the first stage recovers global pattern coherence and the second refines local texture detail, supplemented by a segmentation mask to guide synthesis at ambiguous boundaries. We verify our approach on a collection of self-captured everyday objects and demonstrate globally coherent, materially grounded gigapixel imagery.