Interactive 3D Gaussians exported as .glb
(RealEstate10K).
Drag to orbit, scroll to zoom.
Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manner, overlooking the fundamental distinction between view-independent and view-dependent appearance, which results in suboptimal rendering quality. In this paper, we present AESplat, a novel and general framework for pose-free feed-forward 3DGS that introduces an effective decoupled appearance modeling strategy based on an analysis of SH, enabling higher-quality rendering. Specifically, AESplat directly derives the zeroth-order SH coefficient, which represents the base view-independent appearance component, from the input images without training. The higher-order SH coefficients are subsequently predicted by a shallow multilayer perceptron equipped with two efficient 3D-aware inductive biases to model view-dependent appearance variations. Extensive experiments across multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving a 0.8 dB improvement in PSNR over the pose-free method NAS3R and a 1.1 dB improvement over the pose-required method DepthSplat on the RealEstate10K dataset.
Given unposed images, AESplat jointly predicts camera poses and Gaussian primitives in a single forward pass, with appearance attributes obtained through decoupled appearance modeling. Leveraging the strong observation provided by the input image itself, the zeroth-order SH coefficient of each pixel-aligned Gaussian is directly derived from its corresponding pixel RGB value via Image-to-Direct Component (I2DC), without training. Higher-order SH coefficients are predicted by a view-dependent appearance head that takes two effective 3D-aware inductive biases—Gaussian-to-Views Spatial Relation Embedding (GVSRE) and Warped Color Map (WCM)—enabling more accurate modeling of view-dependent appearance. Built upon a pose-free feed-forward backbone, this decoupled strategy produces higher-quality Gaussian representations and better novel-view rendering.
| Method | RealEstate10K | ACID | ||||
|---|---|---|---|---|---|---|
| PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | |
| Supervised Pose-required | ||||||
| pixelSplat | 23.859 | 0.808 | 0.184 | 25.889 | 0.780 | 0.194 |
| MVSplat | 24.012 | 0.812 | 0.175 | 25.561 | 0.775 | 0.195 |
| DepthSplat | 25.595 | 0.852 | 0.145 | - | - | - |
| YoNoSplat | 24.233 | 0.813 | 0.162 | - | - | - |
| Supervised Pose-free | ||||||
| Splatt3R | 18.688 | 0.337 | 0.596 | 18.060 | 0.510 | 0.407 |
| NoPoSplat | 25.033 | 0.838 | 0.160 | 25.961 | 0.781 | 0.189 |
| Self-Supervised Pose-free | ||||||
| SelfSplat | 19.152 | 0.680 | 0.328 | 22.089 | 0.694 | 0.298 |
| PF3Splat | 21.042 | 0.739 | 0.233 | 21.206 | 0.632 | 0.293 |
| SPFSplat | 25.484 | 0.847 | 0.153 | 26.070 | 0.781 | 0.186 |
| SPFSplatV2-L | 25.668 | 0.855 | 0.137 | 26.674 | 0.806 | 0.162 |
| NAS3R | 25.888 | 0.861 | 0.136 | 26.832 | 0.813 | 0.160 |
| AESplat (Ours) | 26.691 | 0.872 | 0.128 | 27.679 | 0.825 | 0.151 |
Table 1. Performance comparison of novel view synthesis on RealEstate10K and ACID datasets. We report the average metrics across all test scenes. The best and second-best results are highlighted. − indicates that the result was not reported in the original paper.
| Method | ACID | DL3DV | ScanNet++ | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | PSNR↑ | SSIM↑ | LPIPS↓ | |
| Supervised Pose-required | |||||||||
| pixelSplat | 25.477 | 0.770 | 0.207 | 18.688 | 0.582 | 0.354 | 18.422 | 0.720 | 0.278 |
| MVSplat | 25.525 | 0.773 | 0.199 | 17.786 | 0.545 | 0.357 | 17.138 | 0.687 | 0.297 |
| DepthSplat | 26.012 | 0.791 | 0.185 | 19.553 | 0.611 | 0.285 | 20.775 | 0.760 | 0.254 |
| YoNoSplat | 24.246 | 0.721 | 0.222 | 19.636 | 0.594 | 0.311 | 21.075 | 0.744 | 0.254 |
| Supervised Pose-free | |||||||||
| NoPoSplat | 25.764 | 0.776 | 0.199 | 19.974 | 0.612 | 0.305 | 22.136 | 0.798 | 0.232 |
| Self-Supervised Pose-free | |||||||||
| SelfSplat | 22.204 | 0.686 | 0.316 | 15.047 | 0.410 | 0.498 | 13.277 | 0.538 | 0.534 |
| SPFSplat | 25.965 | 0.781 | 0.190 | 19.172 | 0.573 | 0.315 | 19.971 | 0.738 | 0.265 |
| SPFSplatV2-L | 26.361 | 0.796 | 0.169 | 19.743 | 0.613 | 0.277 | 21.796 | 0.811 | 0.200 |
| NAS3R | 26.663 | 0.807 | 0.166 | 19.842 | 0.628 | 0.274 | 21.028 | 0.799 | 0.210 |
| AESplat (Ours) | 27.375 | 0.819 | 0.156 | 20.430 | 0.648 | 0.259 | 21.892 | 0.812 | 0.199 |
Table 2. Cross-dataset generalization. All methods are trained on RealEstate10K and evaluated in a zero-shot setting on ACID, DL3DV, and ScanNet++. Best and second-best are highlighted.
| Method | RealEstate10K | ||
|---|---|---|---|
| PSNR↑ | SSIM↑ | LPIPS↓ | |
| NAS3R-m | 25.814 | 0.856 | 0.149 |
| +Ours | 26.294 | 0.860 | 0.141 |
| SPFSplatV2 | 25.693 | 0.853 | 0.149 |
| +Ours | 26.181 | 0.859 | 0.141 |
Table 3. Baseline Generality. Our approach consistently improves the performance of different baselines.
Qualitative comparison on RealEstate10K. The leftmost column shows the two-view context images.
Qualitative comparison on ACID. The leftmost column shows the two-view context images.
Cross-dataset generalization (trained on RealEstate10K): DL3DV (first four rows), ACID (fifth and sixth rows), and ScanNet++ (last two rows).