Score
Design and implement algorithms that assemble predictions computed on overlapping subregions (patches or subdomains) into a single, seamless image or volume. These methods enforce inter-subdomain continuity across overlap regions, preserve small-scale variability and texture, and mitigate seam artifacts when merging inpainted or reconstructed subdomain outputs.
This study addresses the challenge of artifacts introduced by tiling strategies when generating large-scale images with deep learning, particularly in cryo-electron microscopy data. Using CycleGAN, the authors systematically evaluate three tiling approaches combined with 2D and 3D patching schemes in terms of image generation quality and their impact on downstream mitochondrial segmentation performance. They find that commonly used perceptual metrics such as FID fail to capture subtle artifacts that significantly degrade segmentation accuracy. To mitigate this, the work proposes an orthogonal-direction ensemble prediction strategy that effectively enhances generation quality for low-quality volumetric data. Experimental results show that 3D models yield marginally better results than 2D models under artifact-free tiling but incur substantially higher computational costs; in contrast, 2D models enable larger batch sizes and more stable training. This work highlights a critical disconnect between standard evaluation metrics and task-specific performance in biomedical image domain adaptation.
To address severe seam artifacts caused by global alignment failure in large-disparity image stitching, this paper proposes embedding a Local Patch Alignment Module (LPAM) into the seam cutting pipeline, enabling an integrated “cut-while-align” paradigm. Specifically, image patches are extracted from low-quality seam regions; dense correspondences are established via SIFT flow; and pixel-level realignment is achieved through local affine or homography optimization, followed by multi-frequency seamless blending. This work is the first to tightly couple local alignment with seam cutting, breaking the conventional two-stage “align-then-cut” framework. Evaluated on large-disparity datasets, our method achieves a 2.1 dB PSNR gain and a 0.045 SSIM improvement over baselines, while increasing inference time by less than 8%, thus striking a favorable balance between accuracy and efficiency.
This work addresses the degradation in layout-to-image generation quality caused by dense overlapping bounding boxes. It identifies two core challenges: large-scale spatial overlap and low semantic discriminability among overlapping instances. To mitigate evaluation bias in existing benchmarks (e.g., CLEVR-Layout), which favor low-overlap scenarios, we propose OverLayScore—a quantitative metric for overlap complexity—and introduce OverLayBench, the first high-overlap, balanced benchmark. We conduct the first systematic analysis of overlap’s impact on generation performance and fine-tune CreatiLayout-AM using amodal mask annotations, significantly enhancing modeling of overlapping regions. Experiments demonstrate substantial performance degradation of state-of-the-art methods under high OverLayScore conditions. Our work fills a critical gap in evaluating complex overlap scenarios, establishing a new standard and actionable directions for future research.
Diffusion-based image inpainting and extension often suffer from boundary artifacts and unnatural fusion due to color mismatch and structural discontinuity across the mask boundary. To address this, we propose a two-stage collaborative optimization framework. First, we design a color-bias-correcting variational autoencoder (VAE) that explicitly models and compensates for cross-boundary chromatic shifts. Second, we introduce an appearance-structure disentangled two-stage diffusion training paradigm: pixel-level appearance alignment is enforced in the image space, while geometric consistency is preserved via latent-space constraints. The framework enables end-to-end joint optimization, significantly reducing both pixel-wise L1 error and perceptual discontinuities. Evaluated on benchmarks including Places2 and COCO-Stuff, our method achieves state-of-the-art visual quality—producing seamless, semantically coherent results with imperceptible boundaries and no visible stitching artifacts.
Existing fragment reassembly methods often rely on assumptions of regular geometric shapes, limiting their effectiveness in real-world scenarios—such as archaeology—where fragments exhibit highly irregular geometries. To address this, we propose a hybrid geometric-image compatibility modeling framework that requires no prior assumptions about shape, size, or content. Our approach comprises three key components: (1) a generative model simulating archaeological erosion to produce realistic irregular fragments; (2) a pairwise discriminative mechanism jointly optimizing edge-based geometric alignment and local texture feature matching; and (3) a puzzle-oriented, neighborhood-level evaluation metric. Integrated into an end-to-end archaeological jigsaw solving pipeline, our method achieves state-of-the-art performance on the RePAIR 2D benchmark—improving neighborhood accuracy by +4.2% and recall by +5.8%. It significantly enhances robustness and accuracy in compatibility assessment for geometrically complex fragments.
Existing 3D vision systems rely on spatial and appearance overlap between views, making them ill-suited for reconstruction tasks under non-overlapping viewpoints—such as those encountered in distributed robotic systems or crowdsourced data collection. This work formally defines and addresses, for the first time, the problem of non-overlapping multi-view 3D reconstruction by introducing GLADOS, an architecture-agnostic framework comprising three stages: generative bridging, robust coarse reconstruction, and iterative contextual refinement. GLADOS leverages foundation models to synthesize intermediate views, performs globally aligned coarse reconstruction, and iteratively expands and optimizes geometry with consistency constraints to achieve geometrically complete and semantically coherent reconstructions. Evaluated on a newly curated benchmark dataset, GLADOS significantly outperforms existing methods, effectively mitigating geometric fragmentation and semantic inconsistency, thereby offering the first viable solution for zero-overlap reconstruction scenarios.
This work addresses the limitations of traditional image scaling methods, which often distort semantically important regions. The authors implement a content-aware seam carving algorithm that efficiently computes minimum-energy 8-connected monotonic seams via dynamic programming, supporting image reduction, ordered insertion for enlargement, multi-pass resizing, and user-guided object preservation or removal through masks. The system faithfully reproduces both Avidan and Shamir’s original seam carving approach and Rubinstein’s forward energy criterion, while also providing visualizations of energy maps and carved seams. Experimental results demonstrate that the method effectively preserves salient content across diverse natural images, significantly outperforming conventional scaling techniques and confirming its robustness and practical utility.
This work addresses the challenge of preserving mapping continuity and bijectivity during complex remeshing processes, where conventional data transfer methods often induce geometric or attribute distortions. The authors propose a composite mapping framework based on local bijective atlases, enhanced by a Shared Scaffold structure that guarantees global bijectivity. The approach is generalized to support a variety of remeshing operations and, for the first time, enables the construction of bijective mappings on 3D tetrahedral remeshings by innovatively integrating Steinitz’s theorem with Maxwell–Cremona lifting theory. This framework facilitates precise tracking of geometric entities—including points, curves, and surfaces—across remeshing sequences, significantly improving fidelity in high-precision applications such as texture transfer and volumetric simulation.
This work addresses the challenges of large-scale image semantic editing—namely semantic distortion, misalignment, and visible seam artifacts—which often compromise generation quality and content coherence. The authors propose a training-free, model-agnostic black-box pipeline that, for the first time, enables high-quality editing of large images by leveraging closed-source vision-language models (VLMs) without fine-tuning. The method integrates overlapped tiling, black-box VLM inpainting, geometric and color consistency correction, seam-risk-aware multi-candidate ranking, and dynamic-programming-based curved seam blending. This framework substantially reduces seam visibility, supports arbitrary-region semantic modifications, and is compatible with diverse off-the-shelf VLMs, achieving natural and seamless editing results without requiring model adaptation.