MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale
2026-07-27 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors present MicroZoom, a tool that creates very large, detailed images of objects by combining regular photos with a few close-up microscope shots. Instead of making an exact copy, MicroZoom generates a believable image that shows microscopic textures across the entire object. To do this, it first ensures big patterns like fabric weaves look right, then adds tiny texture details, using masks to handle tricky areas. They tested this on everyday items and produced large images that stay true to the material's look.
gigapixel image synthesissuper-resolutionmicroscopic scaletexture synthesispattern coherenceimage segmentationcascaded designmagnificationmaterial boundariesimage reconstruction
Authors
Huy Huynh, Jingwei Ma, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz
Abstract
We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of consumer-grade microscope close-ups, MicroZoom synthesizes a seamless, gigapixel-resolution image grounded in the material character of the real references, enabling exploratory visualization of microscopic texture across the full spatial extent of an object. Our goal is plausible synthesis, not exact reconstruction. We focus on full-image, reference-based, extreme-scale super-resolution at magnification levels of up to 350x, a setting that introduces two major challenges: (1) recovering texture-specific detail from highly lossy inputs near ambiguous material boundaries, and (2) preserving correct large-scale pattern structure, such as the repeating geometry of a fabric weave, across millions of local predictions. We address these with a two-stage cascaded design, where the first stage recovers global pattern coherence and the second refines local texture detail, supplemented by a segmentation mask to guide synthesis at ambiguous boundaries. We verify our approach on a collection of self-captured everyday objects and demonstrate globally coherent, materially grounded gigapixel imagery.