LightFuse: Relightable Interactive Gaussian Scene Reconstruction
via Multi-Scan Fusion and 2D Gaussian Ray Tracing

Haonan Zhou1, Gaoxiang Linghu1, Youlin Jia2, Hongyu Cui1
Kewei Wei3, Kaiyue Zhou4, Bruce X.B. Yu1, Gaoang Wang1,3,*
1ZJU-UIUC Institute, Zhejiang University
2College of Mathematics, Sichuan University
3College of Computer Science and Technology, Zhejiang University
4Chengdu Minto Tech
* Corresponding author

LightFuse reconstructs interactive scenes from multi-state, multi-illumination observations, enabling novel-state synthesis, relighting, and object material editing.

LightFuse fuses a shared background and movable objects, refines their surfaces for ray tracing, and separates shared material from state-specific illumination.

Abstract

Relightable interactive scene reconstruction aims to build an editable 3D model from scans of different object arrangements and render new layouts under novel illumination. Existing methods either bake lighting into appearance or recover material and illumination only for fixed scenes, leaving edited layouts with inconsistent shadows and indirect lighting. We present LightFuse, a 2D Gaussian framework that extends interactive scene reconstruction with explicit material-illumination decomposition and physically based relighting. LightFuse first fuses observations across states to reconstruct a shared background and movable objects. It then conducts ray-tracing-oriented geometry refinement to produce more complete and consistent surfaces. On the refined geometry, staged training with differentiable one-bounce ray tracing separates shared metallic-roughness material from state-specific environment lighting. The resulting scene supports object rearrangement, material editing, and relighting, while ray tracing recomputes appearance after each interaction. Experiments across synthetic scenes demonstrate state-of-the-art relighting quality, outperforming the strongest baseline by +9.74 dB PSNR and +0.121 SSIM on average.

Methodology

Pipeline of LightFuse

Given multi-state RGB images, masks, normals, and depth, LightFuse first reconstructs an occlusion-complete shared background and movable object components with 2D Gaussian surfels. Cross-state supervision then refines geometry for ray tracing through coverage, scale, depth, and normal regularization. Finally, staged inverse rendering shares material across states while optimizing a separate environment map for each state. Differentiable one-bounce path tracing enables novel-state relighting and material editing with updated light transport.

Results

Novel-State Relighting

Qualitative comparison of novel-state relighting on synthetic scenes

Qualitative comparison under unseen environment maps. LightFuse produces more consistent shadows, reflections, and low-texture surfaces after the scene layout changes.

Real-World Results

Qualitative comparison of novel-state synthesis on real scenes

Novel-state synthesis on captured scenes with changes in object layout and illumination. Red boxes highlight cast shadows and reflections after object rearrangement.

Scene Property Reconstruction

Animated predictions show the recovered intrinsic appearance and geometry throughout the camera trajectory. Together, base color, depth, surface normals, roughness, and metallic values provide the scene properties used for physically based relighting.

Base Color
Depth
Surface Normals
Roughness
Metallic
Base Color
Depth
Surface Normals
Roughness
Metallic

Additional Visual Analysis

Additional held-out-view analyses show that LightFuse recovers cleaner base color and roughness while producing more complete depth and surface normals across Bedroom, Kitchen, Livingroom, and Playroom.

Base-color comparison between the reference, ReCap, ReCap Adapted, and LightFuse
Base Color. LightFuse reduces illumination-dependent artifacts in the recovered material.
Roughness comparison between the reference, ReCap, ReCap Adapted, and LightFuse
Roughness. Shared multi-state material produces more stable roughness estimates.
Depth and normal comparison between the reference, ReCap, ReCap Adapted, and LightFuse
Geometry. Explicit depth and normal supervision yields cleaner boundaries and surface orientations.

Additional Relighting Comparisons

Detailed novel-state relighting comparisons across four scenes and multiple environment maps
Close-up comparisons across all four synthetic scenes. LightFuse better preserves target illumination, shadows, reflections, and low-texture surfaces after object rearrangement.
View all 11 unseen lighting conditions for each scene

BibTeX

@article{zhou2026lightfuse,
  author = {Zhou, Haonan and Linghu, Gaoxiang and Jia, Youlin and Cui, Hongyu and
            Wei, Kewei and Zhou, Kaiyue and Yu, Bruce X. B. and Wang, Gaoang},
  title  = {LightFuse: Relightable Interactive Gaussian Scene Reconstruction
            via Multi-Scan Fusion and 2D Gaussian Ray Tracing},
  year   = {2026},
  eprint = {2608.29269},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV}
}