CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance

arXiv cs.AI Papers

Summary

This paper introduces CoIn, a novel framework for 3D scene inpainting that bridges 2D diffusion models and 3D Gaussian Splatting via a multi-stage consistency pipeline, enabling both object removal and insertion with flexible masks.

arXiv:2606.27584v1 Announce Type: cross Abstract: 3D scene inpainting is essential for reconstructing areas corrupted by occlusions or limited viewpoints. While recent methods leverage Gaussian Splatting (GS) for efficient 3D editing, they often depend on precise multi-view segmentation masks and are inherently constrained to object removal tasks. We propose CoIn, a novel framework that bridges 2D inpainting models and 3DGS through a multi-stage consistency pipeline. Our approach first generates initial inpainted images using a diffusion model, enabling the use of arbitrary-shaped masks and diverse tasks like object insertion. We then introduce Reference Adaptive GS with Feature Attention to reconstruct a coarse 3D scene by adaptively weighing towards a reference view (2D -> 3D). This 3D representation provides geometric guidance to the diffusion process via GS-based Reference Feature Warping, ensuring multi-view consistency (3D -> 2D). Finally, a Texture-Enhancing Discriminator refines the 3D scene to achieve high photometric realism (2D -> 3D). Experiments show that CoIn, effectively leveraging bidirectional information flow, achieves state-of-the-art performance and effectively handles both object removal and object insertion with flexible mask input.
Original Article
View Cached Full Text

Cached at: 06/29/26, 05:29 AM

# CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance
Source: [https://arxiv.org/html/2606.27584](https://arxiv.org/html/2606.27584)
11institutetext:LG Electronics, Seoul, South Korea11email:hana1106\.kim@lge\.com22institutetext:KAIST, Daejeon, South Korea22email:\{minjekim, kimtaekyun\}@kaist\.ac\.kr###### Abstract

3D scene inpainting is essential for reconstructing areas corrupted by occlusions or limited viewpoints\. While recent methods leverage Gaussian Splatting \(GS\) for efficient 3D editing, they often depend on precise multi\-view segmentation masks and are inherently constrained to object removal tasks\. We propose CoIn, a novel framework that bridges 2D inpainting models and 3DGS through a multi\-stage consistency pipeline\. Our approach first generates initial inpainted images using a diffusion model, enabling the use of arbitrary\-shaped masks and diverse tasks like object insertion\. We then introduce Reference Adaptive GS with Feature Attention to reconstruct a coarse 3D scene by adaptively weighing towards a reference view \(2D→\\rightarrow3D\)\. This 3D representation provides geometric guidance to the diffusion process via GS\-based Reference Feature Warping, ensuring multi\-view consistency \(3D→\\rightarrow2D\)\. Finally, a Texture\-Enhancing Discriminator refines the 3D scene to achieve high photometric realism \(2D→\\rightarrow3D\)\. Experiments show that CoIn, effectively leveraging bidirectional information flow, achieves state\-of\-the\-art performance and effectively handles both object removal and object insertion with flexible mask input\.

![[Uncaptioned image]](https://arxiv.org/html/2606.27584v1/x1.png)

Figure 1:Key impact of inconsistent masks on 3D\-first vs\. 2D\-first pipelines\.\(a\) Prior works, which typically adopt a 3D\-first pipeline, often fail under inconsistent 2D segmentation masks across views, leading to excessive removal of regions and incorrect 3D segmentation\. \(b\) In contrast, CoIn utilizes a 2D\-first pipeline that comprehensively integrates 2D and 3D inpainting branches under Gaussian Splatting guidance\. This approach ensures spatial and semantic consistency, enabling the use of arbitrary\-shaped masks and supporting diverse tasks such as object removal and insertion \(c\)\.
## 1Introduction

Inpainting 3D scene is essential for reconstructing incomplete or corrupted scenes arising from occlusions, sensor noise, or limited viewpoints\. Recent advances in generative models have expanded this scope from simple object removal to editing and object insertion\. To achieve this, various methods directly optimize Neural Radiance Fields \(NeRF\)\[mildenhall2020nerf\]using priors such as RGB and depth\[mirzaei2023spin\]or diffusion\-based\[chen2024mvip\]\. However, the implicit representation of NeRF hinders direct geometry manipulation, necessitating complex volumetric optimization for precise removal, insertion, or shape editing\.

To perform more geometric and efficient editing, subsequent methods\[wang2024learning,ye2024gaussian,shi2025imfine,huang20253d,wu2025aurafusion360\]adopt the ‘3D\-first’ pipeline, which denotes the strategy of first reconstructing a 3D Gaussian Splatting \(3DGS\)\[kerbl20233d\]scene from the original images to provide a structural basis before performing inpainting\. However, this sequence requires accurate segmentation masks \(e\.g\., obtained from SAM2\[ravi2024sam\]\) across multiple views to precisely isolate target regions in 3D space\. Furthermore, methods that initiate inpainting by pruning these reconstructed Gaussians, such as IMFine\[shi2025imfine\], 3DGIC\[huang20253d\], and AuraFusion\[wu2025aurafusion360\], are inherently restricted to removal tasks\. Although GScream\[wang2024learning\]avoids initial pruning, it is based only on a single 2D inpainted reference image, which limits its ability to maintain consistency across large viewpoint changes\.

By contrast, the ‘2D\-first’ pipeline denotes a strategy that prioritizes performing 2D inpainting across multiple views, followed by leveraging 3D information, such as optical flow\[cao2024mvinpainter\]or meshes\[barda2025instant3dit\], to ensure consistency among the generated images\. Although pipelines like MVInpainter\[cao2024mvinpainter\]or Instant3dit\[barda2025instant3dit\]benefit from the diverse editing capabilities of 2D generative models, they often rely on fine\-tuning with refined datasets or require additional accurate 3D labels\.

We propose a novel framework that combines the flexibility of 2D\-first pipelines with the efficient optimization of 3DGS\. By integrating these approaches, our method supports arbitrary\-shaped masks and handles diverse inpainting tasks with high efficiency\. Multi\-view consistency is maintained by guiding the 2D inpainter with 3D priors from Reference Adaptive GS, while a Texture Enhancing Discriminator further refines photometric realism\. We demonstrate our approach through both quantitative and qualitative experiments for object removal and insertion tasks\. In summary, our contributions are as follows:

- •We propose CoIn, a novel framework that seamlessly bridges generative 2D synthesis and explicit 3D reconstruction via a multi\-stage pipeline\.
- •We introduce Reference Adaptive Gaussian Splatting with Feature Attention \(Ref\-GS\), which assigns adaptive weights to each view to optimize 3DGS toward the reference viewpoint\.
- •We present Consistency Loss Guidance \(CLG\), which leverages GS\-based Reference Feature Warping to enforce 3D consistency with the reference image during the denoising process\.
- •The Texture Enhancing Discriminator \(TE\-D\) mitigates blurriness in generated patches by learning the distribution of real image patches\.

## 2Related Work

### 2\.12D Image Inpainting

2D inpainting refers to the task of reconstructing missing or masked regions in an image by leveraging contextual cues from visible areas\. Early approaches mainly filled missing areas by copying nearby content, often resulting in visually inconsistent completions\[efros1999texture\]\. With the advent of neural networks trained on datasets, recent methods\[wang2024gridformer,lugmayr2022repaint,ju2024brushnet,saharia2022palette,wang2023towards\]used combinations of convolution networks, such as GAN\[yu2022high\], Fourier convolution\[suvorov2022resolution\], or Wavelet decomposition\[jeevan2023wavepaint\]\.

Meanwhile, diffusion\-based models are widely adopted for 2D inpainting due to their ability to produce diverse yet highly feasible results\[saharia2022palette,lugmayr2022repaint\]\. RePaint\[lugmayr2022repaint\]extends the concept of unconditional DDPMs using the known region of an image as a condition, allowing the denoising process to serve as inpainting for the unknown region\. Recent approaches\[ju2024brushnet,wang2023towards\]incorporate latent diffusion models\[rombach2022high\]to inpainting tasks\. Stable diffusion \(SD\) inpainting pipeline\[rombach2022high\]takes the masked image as input and performs denoising directly in the latent space, achieving faster inference while producing plausible and realistic results\.

However, using pure 2D image inpainters without any additional modification, stochastic sampling in the denoising process leads to severely inconsistent results given small viewpoint changes across multi\-view images, even for the same images\. These inconsistencies are not suitable for subsequent 3D reconstruction, causing inaccurate geometries and blurriness\. While several methods\[deng2023mv,shi2023mvdream,poole2022dreamfusion,tang2024lgm,ai2024dream360,liu2023zero,shen2023anything,barda2025instant3dit,zhuang2024tip,cao2024mvinpainter,weber2024nerfiller,kim2025srhand\]consider 3D consistency for 2D generation and editing, they often exhibit trade\-offs between computational efficiency and structural flexibility\. Specifically, MVInpainter\[cao2024mvinpainter\]and NeRFiller\[weber2024nerfiller\]attempt to enforce consistency via flow supervision or shared grid priors; however, these approaches still encounter limitations in mask shape flexibility, output resolution, or the need for per\-dataset adaptation\.

### 2\.2Guidance for Latent Diffusion Model

There exist various strategies\[kim2024arbitrary,wang2025lldiffusion,ho2022classifier\]to enable diffusion models to perform specific tasks or adapt to particular domains, specifically, fine\-tuning\[hu2022lora\], incorporating a pretrained adapter\[ye2023ip,mou2024t2i\], or injecting guidance\[bansal2023universal,yu2023freedom,song2023loss\]to align the model with a desired task or dataset\. However, fine\-tuning or training adapters generally relies on task\-specific training data, which limits the flexibility of per\-scene adaptation\.

Instead, guidance\-based control methods\[bansal2023universal,yu2023freedom,song2023loss\]introduce additional control signals to the denoising process at inference time\. FreeDoM\[yu2023freedom\]formulates an energy function that measures the distance between the given conditionccand the noisy intermediate resultxtx\_\{t\}\. By minimizing the value of the energy function during denoising steps, diffusion models can generate desirable results without additional training\. Our method leverages guidance\-based control to generate 3D\-consistent inpainted images\. We incorporate guidance from the trained 3DGS within the energy function for the inpainting diffusion model\.

### 2\.33D Scene Inpainting

NeRF\-based 3D inpainting methods\[mirzaei2023spin,chen2024mvip,lin2024taming\]have recently shown strong performance by optimizing an implicit radiance field that can synthesize missing geometry and appearance from multi\-view observations\. However, the implicit formulation typically involves heavy computation and long training/rendering time, and explicit local edits often require re\-optimization or complex volumetric procedures, limiting practical adaptation\.

With 3D Gaussian Splatting \(3DGS\)\[kerbl20233d\], many subsequent approaches have performed inpainting explicitly in 3D space using point\-based scene representations\[ye2024gaussian,shi2025imfine,wang2024learning,huang20253d,wu2025aurafusion360\]\. They usually get one reference image as a guide for inpainting the 3D scene\. However, GScream\[wang2024learning\]relies only on a single reference view that does not cooperate with other viewpoint images, large deviations from the reference view cause 3D inpainting to fail\. In contrast, 3D\-first pipelines remove objects directly in 3D and complete geometry using cues such as Laplacian smoothing to warp the reference image, thereby handling large viewpoint changes\[shi2025imfine,huang20253d,wu2025aurafusion360\]\. Although efficient, they rely on highly accurate 2D segmentation masks to specify a target object consistently across dense views\. As shown in[Fig\.˜1](https://arxiv.org/html/2606.27584#S0.F1), in the absence of 3D consistency in the masks, the segmented regions do not align in the 3D space\. This leads to geometrically inconsistent completion or unintended deletion of non\-target regions and degrades 3D stability\. Moreover, these pipelines are designed primarily for object removal rather than general inpainting tasks, such as insertion\.

To the best of our knowledge, CoIn represents a comprehensive framework that establishes a correlated inpainting pipeline by integrating the generative capabilities of 2D models with the explicit representation of 3D Gaussian Splatting\. Our framework, CoIn, is designed to leverage both 2D and 3D strengths: a 2D inpainting branch enables semantically meaningful edits under arbitrary\-shaped masks, while an explicit 3D inpainting branch ensures strict multi\-view consistency\. Unlike 3D\-first pipelines, CoIn does not require accurate segmentation masks and supports both removal and insertion, while avoiding the cross\-view inconsistency issues common in 2D multi\-view inpainting\.

## 3Preliminary

#### 3D Gaussian Splatting\.

3D Gaussian Splatting \(3DGS\)\[kerbl20233d\]is an explicit point\-based scene representation in which each Gaussian primitive is parameterized by its center positionμ\\mu, rotation matrixRR, scaleSS, colorccand opacityα\\alpha\. The covariance matrix of each Gaussian is defined asΣ=R​S​ST​R⊤\\Sigma=RSS^\{T\}R^\{\\top\}, and Gaussians are represented by

𝒢​\(x\)=exp⁡\(−12​\(x−μ\)⊤​Σ−1​\(x−μ\)\)\\mathcal\{G\}\(x\)=\\exp\\\!\\left\(\-\\tfrac\{1\}\{2\}\(x\-\\mu\)^\{\\top\}\\Sigma^\{\-1\}\(x\-\\mu\)\\right\)\(1\)
Training is conducted by minimizing a combination ofℒ1\\mathcal\{L\}\_\{1\}andℒS​S​I​M\\mathcal\{L\}\_\{SSIM\}between the rendered imageRnR\_\{n\}and the ground\-truth RGB imageInI\_\{n\}:

ℒR​\(Rn,In\)=\(1−λ\)​ℒ1\+λ​\(1−ℒS​S​I​M\)\\mathcal\{L\}\_\{R\}\(R\_\{n\},I\_\{n\}\)=\(1\-\\lambda\)\\mathcal\{L\}\_\{1\}\+\\lambda\(1\-\\mathcal\{L\}\_\{SSIM\}\)\(2\)

#### Scaffold\-GS\.

Unlike the original 3DGS, which directly optimizes the parameters of distributed gaussians, Scaffold\-GS\[lu2024scaffold\]adopts a lightweight anchor\-based structure\. Each anchor has a learnable feature embedding \(anchor feature\), and compact MLPs decode this embedding into neural gaussian attributes within its voxel\-defined local region\. This hierarchical design reduces redundancy by performing densification at the anchor level, achieving comparable rendering quality to the original 3DGS\. The sparse anchor structure also enables efficient point\-based guidance to maintain geometric correspondence across views, while anchor features permit attention to be applied directly in 3D, thereby facilitating 3D inpainting\. We adopt it as our base model for representing the 3D scene, and propose an efficient object handling solution\.

![Refer to caption](https://arxiv.org/html/2606.27584v1/x2.png)Figure 2:Overview of CoIn\. We begin with initial 2D inpainting results and apply Reference GS with Feature Attention to obtain a coarse inpainted 3D scene\. We then use Consistency Loss Guidance for the frozen latent diffusion inpainting model with GS\-based Reference Feature Warping, and the consistency\-preserved results are finally used to fine\-tune the 3D Gaussian Splatting scene𝒢\\mathcal\{G\}with a Texture\-Enhancing Discriminator for realistic details\.

## 4Method

### 4\.1Coherent 2D\-3D Inpainting

#### Overview\.

Our method aims to achieve 3D\-consistent inpainting from a set of object\-present input images\{In\}n=1N\\\{I\_\{n\}\\\}\_\{n=1\}^\{N\}and their corresponding masks\{Mn\}n=1N\\\{M\_\{n\}\\\}\_\{n=1\}^\{N\}\. We begin with initial inpainting results\{I^n\}n=1N\\\{\\hat\{I\}\_\{n\}\\\}\_\{n=1\}^\{N\}obtained from a SD inpainting model, which produces visually plausible completions for each image but fails to achieve consistency across views\. To address this issue, we construct a 3D Gaussian scene from these initial results and refine it to enforce multi\-view coherence while preserving fine details\.

[Fig\.˜2](https://arxiv.org/html/2606.27584#S3.F2)a illustrates our pipeline\. We first apply a coarse stage, the Reference Adaptive GS with Feature Attention up\-weights a selected reference viewI^k\\hat\{I\}\_\{k\}and down\-weights the others via per\-view weights, while regularizing anchor features in the inpainting region by attending to neighboring anchors\. It then enforces multi\-view geometric and appearance consistency by warping reference features on the GS\-derived point cloud and injecting them as guidance into the latent diffusion inpainting model\. With the consistency\-guided inpainted images, we refine the GS scene for high\-frequency textures and photometric realism via an adversarial patch discriminator in a fine stage\.

By integrating these components, our framework produces a 3D inpainted scene that is geometrically consistent, visually coherent, and seamlessly completed across all views\. We introduce the Reference Adaptive GS with Feature Attention \(Ref\-GS\) in[Sec\.˜4\.2](https://arxiv.org/html/2606.27584#S4.SS2), Consistency Loss Guidance \(CLG\) in[Sec\.˜4\.3](https://arxiv.org/html/2606.27584#S4.SS3), and Texture Enhancing Discriminator \(TE\-D\) in[Sec\.˜4\.4](https://arxiv.org/html/2606.27584#S4.SS4)\.

### 4\.2Reference Adaptive GS with Feature Attention

We first construct a 3D Gaussian Splatting \(3DGS\) scene to provide 3D guidance to the 2D inpainting branch\. However, employing vanilla GS often yields view\-dependent inconsistencies and overly blurred renders, which are not suitable for multi\-view inpainting\. To stabilize the scene, we applied a per\-view weight to the rendering loss, encouraging the GS scene to align with the reference viewI^k\\hat\{I\}\_\{k\}while suppressing views that are inconsistent\. Before optimization, we prune points whose projections fall inside the masks, thereby eliminating the influence of the residual object from the COLMAP initialization\[schonberger2016structure\]\.

For each viewnn, we compute a weightWnW\_\{n\}that up\-weights the reference viewI^k\\hat\{I\}\_\{k\}and down\-weights views inconsistent with the reference view:

Wn=\{λr,if​n=k11\+exp⁡\(λr​\(ℒR\(n,t\)−ℒR\(n,t−1\)\)\),if​n≠kW\_\{n\}=\\begin\{cases\}\\lambda\_\{r\},&\\text\{if \}n=k\\\\ \\dfrac\{1\}\{1\+\\exp\(\\lambda\_\{r\}\(\\mathcal\{L\}\_\{R\}^\{\(n,t\)\}\-\\mathcal\{L\}\_\{R\}^\{\(n,t\-1\)\}\)\)\},&\\text\{if \}n\\neq k\\end\{cases\}\(3\)whereℒR\(n,t\)\\mathcal\{L\}\_\{R\}^\{\(n,t\)\}represents the current photometric loss of thenn\-th view oftt\-th iteration,ℒR\(n,t−1\)\\mathcal\{L\}\_\{R\}^\{\(n,t\-1\)\}is from the previous iteration of the same view, andλr\\lambda\_\{r\}is a weight parameter to the reference view\. This fuction leads the GS scene to better reflect the reference image\.

Although GScream\[wang2024learning\]employs cross\-attention to propagate texture from the surrounding region into the inpainted region, it is sensitive to variations in view range since it is applied only to the reference view\. To address this limitation, we use unidirectional attention on anchor features for every training view\. We update the featuresfinAf^\{A\}\_\{\\text\{in\}\}viafinA~=Attn​\(Q=finA,K=fsideA,V=fsideA\)\\tilde\{f^\{A\}\_\{\\text\{in\}\}\}=\\mathrm\{Attn\}\(Q=f^\{A\}\_\{\\text\{in\}\},\\ K=f^\{A\}\_\{\\text\{side\}\},\\ V=f^\{A\}\_\{\\text\{side\}\}\)wherefinAf^\{A\}\_\{\\text\{in\}\}andfsideAf^\{A\}\_\{\\text\{side\}\}denote anchor features inside the inpainting region and in its neighborhood, respectively\. This update enhances the consistency between the inpainted region and the surrounding areas, resulting in more coherent textures\.

In addition, we apply a depth loss in the reference view using the estimated depth, making the GS\-derived point cloudPkP\_\{k\}more stable and better aligned with the reference view geometry\. We begin by using a monocular depth estimator\[yang2024depth\]to estimate depth𝒟k\\mathcal\{D\}\_\{k\}fromI^k\\hat\{I\}\_\{k\}\. With a linear transformation, we transform the rendering depthR𝒟R^\{\\mathcal\{D\}\}toR𝒟¯=A∗R𝒟\+B\\bar\{R^\{\\mathcal\{D\}\}\}=A\*R^\{\\mathcal\{D\}\}\+B\. A and B are obtained by solving a least\-squares problem with respect to𝒟k\\mathcal\{D\}\_\{k\}\[wang2024learning,ke2024repurposing,yu2022monosdf\]\. Then the total GS optimization loss is computed as follows:

ℒ=Wn⋅ℒR​\(Rn,In^\)\+λ𝒟​ℒ𝒟\\mathcal\{L\}=W\_\{n\}\\cdot\\mathcal\{L\}\_\{R\}\(R\_\{n\},\\hat\{I\_\{n\}\}\)\+\\lambda\_\{\\mathcal\{D\}\}\\mathcal\{L\}\_\{\\mathcal\{D\}\}\(4\)ℒ𝒟=1H​W​∑‖R𝒟¯−𝒟k‖1\\mathcal\{L\}\_\{\\mathcal\{D\}\}=\\frac\{1\}\{HW\}\\sum\|\|\\bar\{R^\{\\mathcal\{D\}\}\}\-\\mathcal\{D\}\_\{k\}\|\|\_\{1\}\(5\)whereλ𝒟\\lambda\_\{\\mathcal\{D\}\}denotes the weight for the depth loss, and H and W are height and width of the image, respectively\.

### 4\.3Consistency Loss Guidance

We enforce cross‐view consistency among inpainted images with an energy\-based guidance inspired by FreeDoM\[yu2023freedom\]\. During diffusion denoising at steptt, we update the sample to minimize an energy function built from the GS scene \([Sec\.˜4\.2](https://arxiv.org/html/2606.27584#S4.SS2)\), leveraging both its renderings and the reconstructed point cloudPkP\_\{k\}\. We improve the consistency between the reference imageI^k\\hat\{I\}\_\{k\}and the current view decoded imageX\(n,t\)=D​e​c​\(z\(n,t\)\)X\_\{\(n,t\)\}=Dec\(z\_\{\(n,t\)\}\)from the denoising diffusion latentz\(n,t\)z\_\{\(n,t\)\}by introducing GS\-based Reference Feature Warping \([Fig\.˜2](https://arxiv.org/html/2606.27584#S3.F2)b\)\. Given the point cloudPkP\_\{k\}and the camera parameters, we unproject per\-view features into 3D and reproject them into an intermediate viewpointϕ\\phi\. This viewpoint is established by interpolating between the reference \(𝐏k\\mathbf\{P\}\_\{k\}\) and current \(𝐏n\\mathbf\{P\}\_\{n\}\) camera poses; specifically, we apply linear interpolation for translation,𝐭ϕ=\(1−α\)​𝐭k\+α​𝐭n\\mathbf\{t\}\_\{\\phi\}=\(1\-\\alpha\)\\mathbf\{t\}\_\{k\}\+\\alpha\\mathbf\{t\}\_\{n\}, and Spherical Linear Interpolation \(Slerp\) for rotation,𝐪ϕ=Slerp​\(𝐪k,𝐪n;α\)\\mathbf\{q\}\_\{\\phi\}=\\text\{Slerp\}\(\\mathbf\{q\}\_\{k\},\\mathbf\{q\}\_\{n\};\\alpha\), withα=0\.5\\alpha=0\.5\. From each image, we extract \(i\) high\-frequency features with DINO\[caron2021emerging\]upscaled via Feat\-Up\[fu2024featup\], and \(ii\) appearance \(RGB\) features\. Letfϕ​\(⋅;Pk\)f\_\{\\phi\}\(\\cdot;P\_\{k\}\)denote the warp operator that projects the features to the intermediate viewϕ\\phiwithPkP\_\{k\}\. We then form a 3D consistency loss by applying a cosine similarity term to the high\-frequency features and aL​2L2similarity term to the RGB features:

ℒ3​D=1−c​o​s​\(fϕh​r​\(Ik^\),fϕh​r​\(X\(n,t\)\)\)\+‖fϕr​g​b​\(Ik^\)−fϕr​g​b​\(X\(n,t\)\)‖22\\mathcal\{L\}\_\{3D\}=1\-cos\(f\_\{\\phi\}^\{hr\}\(\\hat\{I\_\{k\}\}\),f\_\{\\phi\}^\{hr\}\(X\_\{\(n,t\)\}\)\)\+\|\|f\_\{\\phi\}^\{rgb\}\(\\hat\{I\_\{k\}\}\)\-f\_\{\\phi\}^\{rgb\}\(X\_\{\(n,t\)\}\)\|\|\_\{2\}^\{2\}\(6\)
This GS\-based Reference Feature Warping provides geometry\-aware guidance at stepttto enforce cross\-view consistency\. To specifically account for non\-overlapping regions between the current view and the reference view, we introduce the point\-cloud support maskmnPm\_\{n\}^\{P\}, which identifies areas where valid 3D correspondences are available\. Furthermore, to compensate for imperfections in the reconstructed point cloudPkP\_\{k\}, we incorporate anL​1L1loss against the GS renderingRnR\_\{n\}as a complementary guidance\.

Consequently, the final energy function is formulated as:

ℰ​\(X\(n,t\)\)=mnP⋅ℒ3​D\+\(1−mnP\)⋅‖X\(n,t\)−Rn‖1\\mathcal\{E\}\(X\_\{\(n,t\)\}\)=m\_\{n\}^\{P\}\\cdot\\mathcal\{L\}\_\{3D\}\+\(1\-m\_\{n\}^\{P\}\)\\cdot\|\|X\_\{\(n,t\)\}\-R\_\{n\}\|\|\_\{1\}\(7\)By minimizing the value ofℰ​\(X\(n,t\)\)\\mathcal\{E\}\(X\_\{\(n,t\)\}\), our framework encourages multi\-view consistency and geometry\-aware inpainting by leveraging correspondences from the GS scene within the mask\.

### 4\.4Texture Enhancing Discriminator

Although[Sec\.˜4\.3](https://arxiv.org/html/2606.27584#S4.SS3)produces per\-view inpainting results, they still exhibit residual artifacts and blurriness due to imperfect geometry and strong guidance\. We therefore introduce a Texture\-Enhancing Discriminator \(TE\-D\) that transfers only texture information from the initial inpainting outputI^n\\hat\{I\}\_\{n\}to the GS renderingsRnR\_\{n\}\. Specifically, TE\-D is trained jointly with the GS model on small image patches, treating patches fromI^n\\hat\{I\}\_\{n\}as real and patches from the current GS renderings as fake\. The GS model \(generator\) is trained adversarially against the discriminator, enhancing high\-frequency texture details while preserving the overall geometric structure\. Operating in local patches\[isola2017image\]near the inpainted region, it emphasizes texture fidelity rather than geometry, helping remove guidance\-induced blur \(see[Fig\.˜6](https://arxiv.org/html/2606.27584#S5.F6)A\(c\)\)\. Simultaneously, the depth loss and the rendering loss used in the GS optimization are computed with the CLG inpainting imagesI~n\\tilde\{I\}\_\{n\}\.

ℒ=ℒR​\(Rn,I~n\)\+λD​ℒ𝒟\+λg​e​n​ℒg​e​n​\(Rn,I^n\)\\mathcal\{L\}=\\mathcal\{L\}\_\{R\}\(R\_\{n\},\\tilde\{I\}\_\{n\}\)\+\\lambda\_\{D\}\\mathcal\{L\}\_\{\\mathcal\{D\}\}\+\\lambda\_\{gen\}\\mathcal\{L\}\_\{gen\}\(R\_\{n\},\\hat\{I\}\_\{n\}\)\(8\)whereλg​e​n\\lambda\_\{gen\}denotes the weight for the adversarial loss\. The combination of adversarial texture matching toI^n\\hat\{I\}\_\{n\}and photometric agreement withI~n\\tilde\{I\}\_\{n\}fine\-tunes the inpainted 3D scene𝒢\\mathcal\{G\}, yielding high fidelity while maintaining the multi\-view consistency established in the previous stage\.

## 5Experiments

### 5\.1Experimental Setup

#### Datasets\.

Following prior works, we perform experiments on the SPIn\-NeRF dataset\[mirzaei2023spin\]\. The dataset contains 10 scenes, each providing 100 calibrated images, 60 object\-present and 40 object\-absent\. For evaluation, we inpaint the object\-present images using the provided segmentation masks or bounding boxes derived from them\. After inpainting, we synthesize novel views at the object\-absent viewpoints to compute the metrics\. We also evaluated on the IMFine dataset\[shi2025imfine\], which comprises 20 scenes grouped into four coverage cases \(90°, 120°, 180°, 360°\) having broader view changes than SPIn\-NeRF\. Each scene contains 125 object\-present and 75 object\-absent images, along with binary masks and calibrated camera parameters\. To better show the effect of the reference image, we evaluated ten scenes with90∘90\\,^\{\\circ\},120∘120\\,^\{\\circ\}and180∘180\\,^\{\\circ\}viewpoint coverage\. The360∘360\\,^\{\\circ\}scenes are shown in the supplementary materials\.

#### Evaluation Metrics\.

We report LPIPS\[zhang2018unreasonable\], PSNR, and FID\[heusel2017gans\]on the full images\. To show the performance fidelity in masked regions more effectively, we also report the masked variants, m\-LPIPS, m\-PSNR, and m\-FID, computed only within the ground\-truth segmentation mask\. For the object insertion task, we follow established practices by calculating theC​L​I​Pd​i​rCLIP\_\{dir\}\(CLIP Text\-Image Directional Similarity\) to assess how well the generated 3D content aligns with the provided text instructions\. Additionally, we conduct a user study to evaluate and compare our method with state\-of\-the\-art baselines\. Specifically, we utilize a comparative voting process to assess four key dimensions: multi\-view consistency, overall visual quality, alignment with text prompts and absence of visual artifacts\. Detailed experimental setups and the specific questions used in the user study are provided in the supplementary material\.

Table 1:Quantitative evaluation on the SPIn\-NeRF dataset\[mirzaei2023spin\]\.Note that m\-LPIPS\[zhang2018unreasonable\]and m\-FID\[heusel2017gans\]represent the LPIPS and FID scores within the ground truth segmentation masks\. BB mask denotes a bounding\-box mask derived from the segmentation mask\. The top three results are highlighted inred,orange, andyellow, respectively\.3D RepresentationMask Typem\-LPIPS \(↓\\downarrow\)LPIPS \(↓\\downarrow\)m\-FID \(↓\\downarrow\)FID \(↓\\downarrow\)NeRFSeg maskSPIn\-NeRF\[mirzaei2023spin\]0\.0530\.31153\.449\.6MVIP\-NeRF\[chen2024mvip\]0\.0500\.31173\.450\.5MALD\-NeRF\[lin2024taming\]\\cellcoloroorange 0\.0310\.30113\.544\.7Gaussian SplattingGaussian Grouping\[ye2024gaussian\]0\.037\\cellcoloryyellow 0\.26132\.544\.93DGIC\[huang20253d\]\\cellcolorrred 0\.028\\cellcoloryyellow 0\.2696\.336\.4GScream\[wang2024learning\]\\cellcoloryyellow 0\.032\\cellcoloryyellow 0\.26\\cellcoloryyellow 86\.1\\cellcoloryyellow 31\.2Ours\\cellcoloryyellow 0\.032\\cellcolorrred 0\.23\\cellcolorrred 80\.4\\cellcoloroorange 28\.9BB mask3DGIC0\.0470\.38168\.0104\.1GScream0\.0430\.29104\.634\.2Ours0\.033\\cellcoloroorange 0\.24\\cellcoloroorange 95\.1\\cellcolorrred 26\.6

![Refer to caption](https://arxiv.org/html/2606.27584v1/x3.png)Figure 3:Qualitative results on the SPIn\-NeRF dataset\[mirzaei2023spin\]\. The first row shows the input images and corresponding masks for each scene, followed by the inpainting results of the baseline methods 3DGIC\[huang20253d\], GScream\[wang2024learning\], and Ours in subsequent rows\. By comparing the four columns on the left with the four columns on the right, we can observe the difference between using the segmentation masks and the bounding\-box masks for inpainting\.
#### Implementation Details\.

All experiments were conducted on a single NVIDIA RTX 4090 \(24GB\) GPU\. We utilize Stable Diffusion 2\.0\[rombach2022high\]for 2D inpainting, processing512×512512\\times 512mask\-centered crops\. The regions outside the mask are combined with the ground\-truth image to preserve originality when training the GS model\. Following prior works\[wang2024learning,shi2025imfine\], we empirically select a single reference viewI^k\\hat\{I\}\_\{k\}from the initial 2D inpainting results to serve as a basis for maintaining consistency across the remaining views\. For memory efficiency during the Consistency Loss Guidance \(CLG\) stage, we employ normalized DINO \(dino16\)\[caron2021emerging\]and FeatUp\[fu2024featup\]features extracted from256×256256\\times 256resized images\. The TE\-D module is trained in two distinct stages: first, we perform 10k iterations of fine\-tuning the generator𝒢\\mathcal\{G\}on the CLG\-inpainted imagesI~n\\tilde\{I\}\_\{n\}, followed by 10k iterations of joint adversarial training\. During this training,64×6464\\times 64patches are extracted via mask\-based probabilistic sampling to focus on the inpainted regions\. Detailed hyperparameter settings, including the loss weightsλD\\lambda\_\{D\},λr\\lambda\_\{r\}, andλg​e​n\\lambda\_\{gen\}are provided in the supplementary material\. The entire pipeline requires approximately 2\.5 hours per scene, which is more efficient than typical learning\-based 3D inpainting approaches\.

### 5\.2Quantitative Results

[Tab\.˜1](https://arxiv.org/html/2606.27584#S5.T1)presents quantitative results on the SPIn\-NeRF dataset\[mirzaei2023spin\]compared to baseline methods\. For the setting using segmentation masks, we use the numbers reported in 3DGIC\[huang20253d\]for SPIn\-NeRF\[mirzaei2023spin\], MVIP\-NeRF\[chen2024mvip\], MALD\-NeRF\[lin2024taming\], Gaussian Grouping\[ye2024gaussian\], and 3DGIC\[huang20253d\]\. For fairness, we reproduce GScream\[wang2024learning\]\(with segmentation and bounding\-box masks\) and 3DGIC \(with bounding\-box masks\) using official implementations\. Since both GScream and 3DGIC require a single reference image, we follow the reference preparation protocol specified in each official implementation\. Our method achieves state\-of\-the\-art performance on most entries under both the segmentation and bounding\-box settings\. Notably, the performance gap between the two types of masks is small for our method, indicating that our method is less sensitive to mask shapes\.

[Tab\.˜2](https://arxiv.org/html/2606.27584#S5.T2)compares our method with 3DGIC\[huang20253d\], GScream\[wang2024learning\], and IMFine\[shi2025imfine\]on the IMFine dataset\. For IMFine, we compute the metrics with officially released result images, while 3DGIC and GScream are evaluated through our reproduction using their official implementations\. Our method outperforms baselines across LPIPS\[zhang2018unreasonable\], PSNR, and FID\[heusel2017gans\]under the segmentation mask setting\.

In[Tab\.˜3](https://arxiv.org/html/2606.27584#S5.T3), we evaluate the object insertion task usingC​L​I​Pd​i​rCLIP\_\{dir\}and a user study\. The results indicate that our method also exhibits certain advantages compared to other methods in terms of both semantic alignment and human preference\. In particular, the total sum of votes in the user study may not be equal to 100% because a ‘None of them’ option was provided to ensure an unbiased evaluation, allowing participants to reject all candidates if none were satisfactory\. Detailed information on the setup of the user study, including the specific questions and the number of participants, is provided in the supplementary material\.

Table 2:Quantitative results on the IMFine dataset\[shi2025imfine\]\.Bold numbers indicate the best performance, and underlined numbers represent the second\-best results\.LPIPS \(↓\\downarrow\)PSNR \(↑\\uparrow\)FID \(↓\\downarrow\)3DGIC\[huang20253d\]0\.338620\.29200\.99GScream\[wang2024learning\]0\.200022\.68112\.91IMFine\[shi2025imfine\]0\.174723\.5957\.53Ours0\.168523\.8853\.37Table 3:Quantitative evaluation of object insertion\.We compare our method against baseline models using theC​L​I​Pd​i​rCLIP\_\{dir\}and a user study across four dimensions\. Values in the user study columns represent the percentage of user preference votes\. Bold numbers indicate the best performance\.C​L​I​Pd​i​rCLIP\_\{dir\}: CLIP Text\-Image Direction Similarity; Align with T\.P\. : Align with Text PromptC​L​I​Pd​i​r↑CLIP\_\{dir\}\\uparrowConsistency \(%\)Visual \(%\)Align with T\.P\. \(%\)w/o Artifacts \(%\)Gaussian Editor\[chen2024gaussianeditor\]0\.022220\.82523\.62513\.90017\.000Infusion\[liu2024infusion\]0\.158931\.95031\.95026\.97530\.550Ours0\.162847\.22544\.42559\.12545\.400

### 5\.3Qualitative Results

In[Fig\.˜3](https://arxiv.org/html/2606.27584#S5.F3), we present comparisons on the SPIn\-NeRF dataset\[mirzaei2023spin\]using both segmentation and bounding\-box masks\. The first row shows the input image and its corresponding mask used for training\. In the second row, we observe that 3DGIC\[huang20253d\]fails to preserve the non\-inpainted regions and generates white artifacts inside the inpainted regions\. For GScream\[wang2024learning\], the generated content does not align with nearby regions \(e\.g\., the wooden bench in the third and fourth columns\)\. In contrast, our method generates high\-fidelity images with multi\-view consistency while preserving the non\-inpainted regions\. Moreover, the inpainting quality is maintained when changing from the segmentation to the bounding\-box settings, indicating that our approach is robust to mask shapes\.

We further compare our method against 3DGIC\[huang20253d\], GScream\[wang2024learning\], and IMFine\[shi2025imfine\]on the IMFine dataset, as shown in[Fig\.˜4](https://arxiv.org/html/2606.27584#S5.F4)\. 3DGIC fails to inpaint removed regions when the geometry is complex, and GScream shows inconsistency when the viewpoint variation is large\. Especially in second scene, while IMFine produces high\-fidelity images, residual object traces remain after inpainting\.

Beyond removal, we also perform object insertion across several scenes\. As shown in[Fig\.˜5](https://arxiv.org/html/2606.27584#S5.F5), Gaussian Editor\[chen2024gaussianeditor\]fails to synthesize new objects, merely altering colors without handling the required geometric changes, especially in the“Jansport"scene\. Infusion\[liu2024infusion\]produces a distorted geometry for the apple in the“Apples2"scene and leaves artifacts behind the melon in the“Jansport"scene\. By contrast, our approach synthesizes high\-fidelity objects that are seamlessly integrated into the scene while maintaining strict multi\-view consistency across varying viewpoints\.

![Refer to caption](https://arxiv.org/html/2606.27584v1/x4.png)Figure 4:Qualitative results on the IMFine dataset\[shi2025imfine\]\. Same as in[Fig\.˜3](https://arxiv.org/html/2606.27584#S5.F3), the first row shows the input images and corresponding masks\. We compare the rendering results with 3DGIC\[huang20253d\], GScream\[wang2024learning\], IMFine\[shi2025imfine\]\. Red box in the input image indicates the reference image for 3DGIC, GScream and Ours\.![Refer to caption](https://arxiv.org/html/2606.27584v1/x5.png)Figure 5:Qualitative results for object insertion task\.We perform object insertion tasks on the“Apples2"and“Jansport"scenes of IMFine dataset\[shi2025imfine\]\. We compare the rendering results with Gaussian Editor\[chen2024gaussianeditor\]and Infusion\[liu2024infusion\]\.Table 4:Quantitative results for ablation studies on the SPIn\-NeRF dataset\[mirzaei2023spin\]\.The full model with all key components achieves the best performance\.Methods \(↓\\downarrow\)m\-LPIPSLPIPSm\-FIDFIDw/o Ref\-GS0\.03590\.253487\.0530\.47w/o Adaptive Weight0\.03560\.257182\.0730\.76w/o Feature Attention0\.03550\.255988\.1531\.78w/o CLG0\.04280\.305797\.0436\.47w/o TE\-D0\.05350\.2902132\.5539\.31Full model \(Ours\)0\.03170\.234980\.4428\.92
### 5\.4Ablation Study

#### Effectiveness of key components\.

We conduct ablation studies on the SPIn\-NeRF dataset\[mirzaei2023spin\]to evaluate the effectiveness of our three key components: Reference Adaptive Gaussian Splatting with Feature Attention \(Ref\-GS\), Consistency Loss Guidance \(CLG\), and Texture\-Enhancing Discriminator \(TE\-D\)\. As shown in[Tab\.˜4](https://arxiv.org/html/2606.27584#S5.T4), removing any component leads to a noticeable decrease in performance in all metrics, indicating that each module is essential for high\-quality, consistent inpainting\. Specifically, we further examine the individual components of Ref\-GS by separately ablating Adaptive Weight and Feature Attention\. Excluding either component leads to performance degradation, confirming that both components are vital for aligning the 3DGS scene toward the reference view while maintaining overall structure\. Specifically, removing TE\-D significantly decreases accuracies, especially in the m\-FID/FID\[heusel2017gans\], which indicates that TE\-D improves rendering quality with increasing the photometric reality\.

[Fig\.˜6](https://arxiv.org/html/2606.27584#S5.F6)A shows the qualitative comparisons of the ablation study\. Rows \(a\)–\(d\) correspond to without Ref\-GS, without CLG, without TE\-D, and the Full model, respectively\. Column \(a\) fails to generate cross\-view consistent results \(see the yellow circle\), and column \(b\) shows irregular artifacts\. Although column \(c\) achieves consistent results, it produces overly smooth textures compared to the full model\. This over\-smoothing behavior corresponds to the lowest quantitative scores on all metrics except LPIPS\[zhang2018unreasonable\]\. For LPIPS, the jittering textures observed in the results without CLG lead to even worse scores\.

#### Effectiveness of consistent 2D images on 3D geometry\.

To assess the effect of enforcing 2D consistency during inpainting on the recovered 3D geometry, we present RGB renderings and depth maps in[Fig\.˜6](https://arxiv.org/html/2606.27584#S5.F6)B\. The figure reports 3DGS results obtained by training with three types of inputs: occluded 2D images, 2D inpainting outputs without guidance, and our guided images\. In the first column, the occluded regions remain visible in both the RGB rendering and the depth map\. Although training with unguided 2D inpainting outputs appears to fill the occlusions in the RGB renderings, these renderings remain blurry and the depth maps lack geometric stability\. In contrast, ours shows more plausible and consistent RGB rendering and depth map\. The results indicate that training 3DGS with consistent 2D images improves not only photometric fidelity but also underlying geometric quality\.

![Refer to caption](https://arxiv.org/html/2606.27584v1/x6.png)Figure 6:Results of ablation studies\(A\) Qualitative results for ablation studies on the“Book"scene of SPIn\-NeRF dataset\[mirzaei2023spin\]\. We mark inpainted regions with red boxes and especially highlight inconsistencies with yellow circles\. Rows \(a\)–\(d\) show w/o Ref\-GS, w/o CLG, w/o TE\-D, and the full model, respectively\. \(B\) Depth map comparison on the“Dabao"scene from the IMFine dataset\[shi2025imfine\]\. We highlight the inpainted regions with red boxes\. For each row, we present the RGB rendering and the depth map from the same viewpoint\.

## 6Conclusions

In this paper, we present CoIn, a comprehensive 2D–3D inpainting framework\. By combining the strengths of 2D inpainting and 3D inpainting, our method supports arbitrary\-shaped masks while maintaining cross\-view geometric and appearance consistency\. We achieve this by optimizing the GS scene toward a reference view with adaptive weights \(Ref\-GS\) and using it to guide the 2D inpainter via GS\-based Reference Feature Warping\. With the guided inpainting images and Texture\-Enhancing Discriminator, we fine\-tune the inpainted GS scene to obtain multi\-view consistency and recover high\-frequency textures\. In the experiments, we achieve state\-of\-the\-art performance in both quantitative and qualitative aspects\. Extensive experiments show that CoIn handles multiple inpainting tasks \(e\.g\., removal and insertion\) under both segmentation and bounding\-box masks, outperforming existing 3D inpainting approaches\.

## Acknowledgements

This work was supported by NST grant \(CRC 21015, MSIT\), IITP grant \(RS\-2023\-00228996, RS\-2024\-00459749, RS\-2025\-25443318, RS\-2025\-25441313, RS\-2026\-25526850, RS\-2026\-25522885, MSIT\), KOCCA grant \(RS\-2024\-00442308, MCST\) and InnoCORE program \(N10260110, MSIT\)\.

## References

## ADataset Details

### A\.1SPIn\-NeRF Dataset

For all experiments on the SPIn\-NeRF dataset\[[S1](https://arxiv.org/html/2606.27584#biba.bibx1)\], we use all 10 scenes provided in the benchmark\. Following prior works, we use the resolution of1008×5671008\\times 567, released with the dataset\. The dataset provides calibrated camera parameters in a simple radial distortion model; therefore, we convert them into a pinhole camera model using the COLMAP\[[S5](https://arxiv.org/html/2606.27584#biba.bibx5)\]converter for training 3d Gaussian splatting, while maintaining the original intrinsic and extrinsic parameters\. Enclosed bounding\-box masks are created from segmentation masks provided by the dataset\. These bounding\-box masks are directly used as the inpainting regions in our method and also baseline methods for fair comparisons across all scenes\.

### A\.2IMFine Dataset

We split the IMFine dataset\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]into two subsets based on the range of viewpoints\. According to the authors of IMFine, the 20 scenes are composed of 2 scenes with90∘90^\{\\circ\}coverage, 4 with120∘120^\{\\circ\}, 4 with180∘180^\{\\circ\}, and 10 with360∘360^\{\\circ\}coverage\. The viewpoint ranges of180∘180^\{\\circ\}or less help to demonstrate plausibility with respect to the reference view\. While we present results on the 10 scenes \(except for360∘360\\,^\{\\circ\}\) in the main paper Table 2, we additionally perform experiments on the scenes of360∘360^\{\\circ\}in[Sec\.˜D\.2](https://arxiv.org/html/2606.27584#S4.SS2a)\.

## BAdditional Implementation Details

Regarding the loss weights, we useλr∈\{50,100\}\\lambda\_\{r\}\\in\\\{50,100\\\}depending on the scene, withλD=10\.0\\lambda\_\{D\}=10\.0andλg​e​n=0\.01\\lambda\_\{gen\}=0\.01\. For the diffusion sampling process, the initial inpainting stage is performed with 50 DDIM steps and a ‘uniform’ discretization schedule\. Consistency Loss Guidance \(CLG\) inpainting stage uses 200 steps and a ‘quad’ schedule\. Detailed configuration files and source code will be publicly released to facilitate reproducibility\.

## CUser Study Protocol

To evaluate the perceptual performance of our method, we conducted a user study involving 18 participants\. The evaluation was performed on 4 distinct scenes—covering both the SPIn\-NeRF\[[S1](https://arxiv.org/html/2606.27584#biba.bibx1)\]and IMFine datasets\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]—including the representative examples presented in the main paper\. For each scene, participants were asked to evaluate the results based on the following four dimensions:

- •Consistency: Which model exhibits the best multi\-view consistency for the inserted object?
- •Overall visual quality: Which model produces the highest visual quality? \(Considering artifacts and color naturalness\)
- •Alignment with text prompts: Which result best aligns with the provided text prompt?
- •Absence of visual artifacts: Which model best preserves the original background without adding any unwanted elements or artifacts? \(Compared with the input image\)

The study included results from Gaussian Editor\[[S10](https://arxiv.org/html/2606.27584#biba.bibx10)\], Infusion\[[S11](https://arxiv.org/html/2606.27584#biba.bibx11)\], and our method, along with a ‘None of them’ option\. To ensure an unbiased assessment, the display order of the models was randomized for each question\.

## DAdditional Experiments

We conduct additional experiments on the IMFine datasets\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]with the bounding box masks and incorporating the 3D\-first method into our pipeline\.

### D\.1Bounding Box Masks

To demonstrate the robustness of our method to mask shapes, we present the inpainting results using the bounding\-box masks\. The same scenes as in the main paper Table 2 are used in the experiments\.

#### Quantitative Results\.

[Tab\.˜2](https://arxiv.org/html/2606.27584#S4.T2)shows the quantitative results for the bounding\-box mask setting on the IMFine dataset\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]\. Our method with the bounding\-box masks achieves performance comparable to that of the segmentation\-mask setting, demonstrating its robustness to arbitrary\-shaped masks\.

Table 1:Bounding Box experiments on the IMFine dataset\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]\.We use the same scenes from the main paper, showing robustness with comparable results\. Seg Mask and BB Mask denote the segmentation mask and its derived bounding\-box mask, respectively\.Ours\(Seg Mask\)Ours\(BB Mask\)LPIPS \(↓\\downarrow\)0\.16850\.1759PSNR \(↑\\uparrow\)23\.8823\.67FID \(↓\\downarrow\)53\.3761\.38
Table 2:Quantitative results on360∘360^\{\\circ\}scenes from the IMFine dataset\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]\.We compare our 2D\-first and 3D\-first results for 3D object removal\. Bold numbers indicate the best performance\.Ours\(2D\-First\)Ours\(3D\-First\)LPIPS \(↓\\downarrow\)0\.14630\.1451PSNR \(↑\\uparrow\)22\.3322\.67FID \(↓\\downarrow\)53\.0248\.67

#### Qualitative Results\.

In[Fig\.˜5](https://arxiv.org/html/2606.27584#S6.F5), the last row shows qualitative results using the bounding\-box masks\. 3DGIC\[[S4](https://arxiv.org/html/2606.27584#biba.bibx4)\]and GScream\[[S3](https://arxiv.org/html/2606.27584#biba.bibx3)\]produce noticeable artifacts around the object regions, as these 3D\-first pipelines rely on accurate segmentation masks to remove objects in the 3D scene\. These similar artifacts are also observed in the bottom two rows of[Fig\.˜4](https://arxiv.org/html/2606.27584#S6.F4)\. In contrast, our method maintains stable and coherent inpainting results even with the coarse bounding\-box masks, highlighting its benefits where exact object masks are hard to obtain\.

### D\.2Incorporating 3D\-first Pipeline

Our method mainly follows a 2D\-first pipeline, which shows strength in inpainting with arbitrary\-shaped masks and enabling diverse object insertion tasks\. However, it receives larger inpainting regions \(see the left column of[Fig\.˜2](https://arxiv.org/html/2606.27584#S6.F2)\), as inpainting is performed on 2D images\. Meanwhile, the 3D\-first pipeline offers the advantage of inpainting only the truly occluded or missing regions observable across multiple views \(see the right column of[Fig\.˜2](https://arxiv.org/html/2606.27584#S6.F2)\), when certain conditions are met\. \(i\) The mask should be precise to the object in the image, and show effectiveness \(ii\) when the scene is captured over a wide viewpoint range\. Since our method adopts a hybrid 2D–3D design, it can incorporate the benefits of the 3D\-first pipeline under certain conditions, allowing us to combine the flexibility of 2D inpainting with the 3D\-first pipeline\.

Specifically, we reconstruct a 3DGS scene and remove Gaussian anchors within the masked regions to obtain𝒢r​m​v\\mathcal\{G\}^\{rmv\}\. By rendering𝒢r​m​v\\mathcal\{G\}^\{rmv\}with fixed Gaussian scales, we derive refined 2D masksMnr​m​vM^\{rmv\}\_\{n\}that represent only the truly occluded regions, along with the corresponding rendered imagesRnr​m​vR\_\{n\}^\{rmv\}\. This process is applied to360∘360^\{\\circ\}scenes \(“msi", “desk3", “rocks", “bin"\) in the IMFine dataset\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]\.

For the360∘360^\{\\circ\}scenes, applying our method directly results in significantly larger inpainting regions compared to the other scenes, making it difficult to achieve competitive performance against 3D\-first pipelines such as 3DGIC and IMFine\. Therefore, for the object\-removal setting in these scenes, we also adopt a mask\-based deletion strategy similar to previous 3D\-first approaches\.

#### Quantitative Results\.

[Tab\.˜2](https://arxiv.org/html/2606.27584#S4.T2)shows that incorporating the 3D\-first pipeline into our framework leads to a noticeable performance gain in all evaluation metrics\. By focusing the inpainting task on a reduced, more precise region, our 3D\-first design achieves higher reconstruction quality compared to the 2D\-first alternative\.

#### Qualitative Results\.

[Fig\.˜2](https://arxiv.org/html/2606.27584#S6.F2)provides a qualitative comparison between the 2D\-first and 3D\-first pipeline when integrated into our framework\. The 3D\-first pipeline yields a refined removal maskMnr​m​vM^\{rmv\}\_\{n\}that is substantially smaller than the one produced by the 2D\-first pipeline, as it identifies only the truly occluded regions observed across multiple views\. This reduced masked area leads to noticeably more faithful inpainting, resulting in a closer match to the ground\-truth images \(e\.g\., wrinkles on the tablecloth\)\. Overall, the qualitative results highlight the benefit of incorporating a 3D\-first stage when accurate multi\-view cues are available\.

### D\.3Robustness to Arbitrary Masks

As seen in[Fig\.˜3](https://arxiv.org/html/2606.27584#S6.F3)\(a\), 2D masks obtained by segmentation model often lack 3D consistency \(e\.g\. including non\-target object mask\)\. Our method, however, remains robust to mask qualities, yielding high\-quality results even with irregular masks\. As shown in[Fig\.˜3](https://arxiv.org/html/2606.27584#S6.F3)\(b\), they are obtained by appending random scribbles to object masks\.[Tab\.˜3](https://arxiv.org/html/2606.27584#S4.T3)validates that our method performs comparable to using precise masks, reducing manual efforts, and enhancing utilities\.

Table 3:Performance ofours across different mask typeson SPin\-NeRF dataset\[[S1](https://arxiv.org/html/2606.27584#biba.bibx1)\]Scene \#10\. Best in bold\.Mask Typem\-LPIPS \(↓\\downarrow\)LPIPS \(↓\\downarrow\)m\-FID \(↓\\downarrow\)FID \(↓\\downarrow\)Irregular mask0\.01090\.216364\.6819\.88BB mask0\.01110\.217961\.2217\.93Seg mask0\.01060\.210865\.6417\.28
### D\.4Robustness to Reference View Selection

As a practical guideline, the reference view is selected based on the desired inpainting result \(e\.g\., the most representative angle of the target object\)\. While this choice is flexible, our method is not sensitive to specific viewpoints; as demonstrated in[Tab\.˜4](https://arxiv.org/html/2606.27584#S4.T4), the performance variation across different reference views is small\.

Table 4:Robustness to reference view selection\.We report the mean and standard deviation across three different reference views \(one as used in the main paper and two randomly chosen\) on the SPin\-NeRF dataset\[[S1](https://arxiv.org/html/2606.27584#biba.bibx1)\]Scene \#10\.Metric \(↓\\downarrow\)m\-LPIPSLPIPSm\-FIDFIDResults0\.0107±0\.00010\.0107\\pm 0\.00010\.213±0\.0050\.213\\pm 0\.00560\.4±3\.760\.4\\pm 3\.718\.4±0\.918\.4\\pm 0\.9

## EAdditional Qualitative Results

We present several qualitative results on the SPIn\-NeRF dataset\[[S1](https://arxiv.org/html/2606.27584#biba.bibx1)\]in[Fig\.˜4](https://arxiv.org/html/2606.27584#S6.F4)and IMFine dataset\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]in[Fig\.˜5](https://arxiv.org/html/2606.27584#S6.F5)\. The first column shows the input image and mask\. For each paired row, we show different views of the same scene\. We present further object insertion experiments in[Fig\.˜2](https://arxiv.org/html/2606.27584#S6.F2), covering a diverse set of scenes from both datasets\. These results highlight our method’s capability to perform insertion tasks that remain robust across large viewpoint ranges\.

## FLimitation and Future Works

Our method relies on a 2D inpainting model based on Stable Diffusion \(SD\)\[[S6](https://arxiv.org/html/2606.27584#biba.bibx6)\]2\.0, and therefore the overall performance is constrained by the base model\. Although SD\-based models accept text prompts, their outputs often depend on the mask shapes, sometimes resulting in reference images that are suboptimal for guiding multi\-view consistency\. In addition, when the provided mask does not fully cover the target object, the quality of removing the object degrades in the 2D inpainting stage\. Replacing the base inpainting model with a more advanced architecture such as SD\-XL\[[S9](https://arxiv.org/html/2606.27584#biba.bibx9)\]could further improve the robustness and visual quality of our pipeline\.

### F\.1Failure Analysis

As shown in[Fig\.˜3](https://arxiv.org/html/2606.27584#S6.F3)\(c\), large masks with view changes beyond180∘180^\{\\circ\}show failure cases due to insufficient context for the 2D inpainting model\. However, as discussed in[D\.2](https://arxiv.org/html/2606.27584#S4.SS2a), incorporating a 3D\-first pipeline addresses this limitation\.

![Refer to caption](https://arxiv.org/html/2606.27584v1/x7.png)Figure 1:Qualitative comparison of 2D\-first and 3D\-first pipelines\.The 3D\-first pipeline simplifies inpainting by targeting only truly occluded areas, unlike the more challenging 2D\-first approach with its larger mask\.
![Refer to caption](https://arxiv.org/html/2606.27584v1/x8.png)Figure 2:Qualitative results of object insertion\.The first row presents the input images, masks, and the corresponding text prompts\.

![Refer to caption](https://arxiv.org/html/2606.27584v1/x9.png)Figure 3:Qualitative results\.\(a\) Input views and 3D\-inconsistent masks from 2D segmentation\. \(b\) Irregular masks with original inputs \(left\) and our consistent novel views \(right\)\. \(c\) Failure cases: large masks with view changes beyond180∘180^\{\\circ\}\(left\) lead to blurry textures \(right\)\.![Refer to caption](https://arxiv.org/html/2606.27584v1/x10.png)Figure 4:Additional qualitative results on the SPIn\-NeRF dataset\[[S1](https://arxiv.org/html/2606.27584#biba.bibx1)\]\.![Refer to caption](https://arxiv.org/html/2606.27584v1/x11.png)Figure 5:Additional qualitative results on the IMFine dataset\[[S2](https://arxiv.org/html/2606.27584#biba.bibx2)\]\.

## References

- \[S1\]Ashkan Mirzaei, Tristan Aumentado\-Armstrong, Konstantinos G Derpanis, Jonathan Kelly, Marcus A Brubaker, Gilitschenski, and Alex Levinshtein\. Spin\-nerf: Multiview segmentation and perceptual inpainting with neural radiance fields\. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20669–20679, 2023
- \[S2\]Zhihao Shi, Dong Huo, Yuhongze Zhou, Yan Min, Juwei Lu, and Xinxin Zuo\. Imfine: 3d inpainting via geometry\-guided multi\-view refinement\. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 26694\-26703, 2025\.
- \[S3\]Yuxin Wang, Qianyi Wu, Guofeng Zhang, and Dan Xu\. Learning 3d geometry and feature consistent gaussian splatting for object removal\. In European Conference on Computer Vision, pages 1–17\. Springer, 2024\.
- \[S4\]Sheng\-Yu Huang, Zi\-Ting Chou, and Yu\-Chiang Frank Wang\. 3d gaussian inpainting with depth\-guided cross\-view consistency\. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 26704–26713, 2025\.
- \[S5\]Johannes L Schonberger and Jan\-Michael Frahm\. Structure\-from\-motion revisited\. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016\.
- \[S6\]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer\. High\-resolution image synthesis with latent diffusion models\. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022\.
- \[S7\]Mingqiao Ye, Martin Danelljan, Fisher Yu, and Lei Ke\. Gaussian grouping: Segment and edit anything in 3d scenes\. In European conference on computer vision, pages 162–179\. Springer, 2024\.
- \[S8\]Chenjie Cao, Chaohui Yu, Fan Wang, Xiangyang Xue, and Yanwei Fu\. Mvinpainter: Learning multi\-view consistent inpainting to bridge 2d and 3d editing\. arXiv preprint arXiv:2408\.08000, 2024\.
- \[S9\]Podell, Dustin, et al\. "Sdxl: Improving latent diffusion models for high\-resolution image synthesis\." arXiv preprint arXiv:2307\.01952 \(2023\)\.
- \[S10\]Chen, Yiwen, et al\. "Gaussianeditor: Swift and controllable 3d editing with gaussian splatting\." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition\. 2024\.
- \[S11\]Liu, Zhiheng, et al\. "Infusion: Inpainting 3d gaussians via learning depth completion from diffusion prior\." arXiv preprint arXiv:2404\.11613 \(2024\)\.

Similar Articles

InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

Hugging Face Daily Papers

InfiniSplat presents a feed-forward single-image 3D Gaussian Splatting framework that uses geometry-guided sampling and query-conditioned implicit decoding to achieve surface-aligned Gaussian representation, improving large-baseline monocular view synthesis and generalizing from synthetic indoor training to open-world scenes.

GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens

Hugging Face Daily Papers

GlobalSplat introduces an efficient feed-forward framework for 3D Gaussian splatting that achieves compact and consistent scene reconstruction using global scene tokens, reducing computational overhead and inference time to under 78ms. The method uses a coarse-to-fine training approach to prevent representation bloat while maintaining competitive novel-view synthesis performance with significantly fewer Gaussians (16K) compared to dense baselines.

In-Context Inpainting for Time Series Forecasting

arXiv cs.AI

ICI-Time is a novel framework that reframes time series forecasting as a visual inpainting task, leveraging large vision models to enable adaptable forecasting without fine-tuning or architectural changes.

Painting with Gaussians

Lobsters Hottest

The author describes building an interactive painting tool that uses edge information from an image to guide 2D Gaussian splats as brush strokes, avoiding slow gradient descent methods and producing painting-like results.