AESplat: Advancing Pose-free Feed-forward 3D Gaussian Splatting via Decoupled Appearance Modeling

Under Review

TL;DR

  • AESplat shifts from a unified appearance modeling paradigm to a decoupled one for pose-free feed-forward 3D Gaussian Splatting, substantially improving rendering quality.
  • We propose an efficient decoupled appearance modeling strategy: the zeroth-order SH coefficient (view-independent appearance) is derived directly from input images without training (I2DC), while higher-order SH coefficients (view-dependent appearance) are predicted by a shallow MLP with two 3D-aware inductive biases: GVSRE and WCM.
  • AESplat achieves state-of-the-art novel view synthesis, with a 0.8 dB PSNR gain over the pose-free method NAS3R and a 1.1 dB gain over the pose-required method DepthSplat on RealEstate10K.

Demo

3D Reconstruction

Interactive 3D Gaussians exported as .glb (RealEstate10K). Drag to orbit, scroll to zoom.

Abstract

Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manner, overlooking the fundamental distinction between view-independent and view-dependent appearance, which results in suboptimal rendering quality. In this paper, we present AESplat, a novel and general framework for pose-free feed-forward 3DGS that introduces an effective decoupled appearance modeling strategy based on an analysis of SH, enabling higher-quality rendering. Specifically, AESplat directly derives the zeroth-order SH coefficient, which represents the base view-independent appearance component, from the input images without training. The higher-order SH coefficients are subsequently predicted by a shallow multilayer perceptron equipped with two efficient 3D-aware inductive biases to model view-dependent appearance variations. Extensive experiments across multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving a 0.8 dB improvement in PSNR over the pose-free method NAS3R and a 1.1 dB improvement over the pose-required method DepthSplat on the RealEstate10K dataset.

Method

AESplat pipeline: decoupled appearance modeling with I2DC, GVSRE, and WCM

Given unposed images, AESplat jointly predicts camera poses and Gaussian primitives in a single forward pass, with appearance attributes obtained through decoupled appearance modeling. Leveraging the strong observation provided by the input image itself, the zeroth-order SH coefficient of each pixel-aligned Gaussian is directly derived from its corresponding pixel RGB value via Image-to-Direct Component (I2DC), without training. Higher-order SH coefficients are predicted by a view-dependent appearance head that takes two effective 3D-aware inductive biases—Gaussian-to-Views Spatial Relation Embedding (GVSRE) and Warped Color Map (WCM)—enabling more accurate modeling of view-dependent appearance. Built upon a pose-free feed-forward backbone, this decoupled strategy produces higher-quality Gaussian representations and better novel-view rendering.

Quantitative Results

Method RealEstate10K ACID
PSNR↑SSIM↑LPIPS↓ PSNR↑SSIM↑LPIPS↓
Supervised Pose-required
pixelSplat 23.8590.8080.184 25.8890.7800.194
MVSplat 24.0120.8120.175 25.5610.7750.195
DepthSplat 25.5950.8520.145 ---
YoNoSplat 24.2330.8130.162 ---
Supervised Pose-free
Splatt3R 18.6880.3370.596 18.0600.5100.407
NoPoSplat 25.0330.8380.160 25.9610.7810.189
Self-Supervised Pose-free
SelfSplat 19.1520.6800.328 22.0890.6940.298
PF3Splat 21.0420.7390.233 21.2060.6320.293
SPFSplat 25.4840.8470.153 26.0700.7810.186
SPFSplatV2-L 25.6680.8550.137 26.6740.8060.162
NAS3R 25.8880.8610.136 26.8320.8130.160
AESplat (Ours) 26.6910.8720.128 27.6790.8250.151

Table 1. Performance comparison of novel view synthesis on RealEstate10K and ACID datasets. We report the average metrics across all test scenes. The best and second-best results are highlighted. − indicates that the result was not reported in the original paper.

Method ACID DL3DV ScanNet++
PSNR↑SSIM↑LPIPS↓ PSNR↑SSIM↑LPIPS↓ PSNR↑SSIM↑LPIPS↓
Supervised Pose-required
pixelSplat 25.4770.7700.207 18.6880.5820.354 18.4220.7200.278
MVSplat 25.5250.7730.199 17.7860.5450.357 17.1380.6870.297
DepthSplat 26.0120.7910.185 19.5530.6110.285 20.7750.7600.254
YoNoSplat 24.2460.7210.222 19.6360.5940.311 21.0750.7440.254
Supervised Pose-free
NoPoSplat 25.7640.7760.199 19.9740.6120.305 22.1360.7980.232
Self-Supervised Pose-free
SelfSplat 22.2040.6860.316 15.0470.4100.498 13.2770.5380.534
SPFSplat 25.9650.7810.190 19.1720.5730.315 19.9710.7380.265
SPFSplatV2-L 26.3610.7960.169 19.7430.6130.277 21.7960.8110.200
NAS3R 26.6630.8070.166 19.8420.6280.274 21.0280.7990.210
AESplat (Ours) 27.3750.8190.156 20.4300.6480.259 21.8920.8120.199

Table 2. Cross-dataset generalization. All methods are trained on RealEstate10K and evaluated in a zero-shot setting on ACID, DL3DV, and ScanNet++. Best and second-best are highlighted.

Method RealEstate10K
PSNR↑SSIM↑LPIPS↓
NAS3R-m 25.8140.8560.149
+Ours 26.2940.8600.141
SPFSplatV2 25.6930.8530.149
+Ours 26.1810.8590.141

Table 3. Baseline Generality. Our approach consistently improves the performance of different baselines.

Qualitative Results

Qualitative comparison on RealEstate10K

Qualitative comparison on RealEstate10K. The leftmost column shows the two-view context images.

Qualitative comparison on ACID

Qualitative comparison on ACID. The leftmost column shows the two-view context images.

Cross-dataset qualitative comparison

Cross-dataset generalization (trained on RealEstate10K): DL3DV (first four rows), ACID (fifth and sixth rows), and ScanNet++ (last two rows).