Submitted:
24 August 2026
Posted:
26 August 2026
You are already at the latest version
Abstract
Monocular Gaussian SLAM must recover camera motion, surface structure, and appearance from an RGB sequence without metric depth input or benchmark geometry during reconstruction. This setting is challenged by scale-ambiguous predictions, spatially varying reliability, and the tendency of an unconstrained Gaussian map to absorb pose and depth errors into its geometric and appearance parameters. We present Mono3DGS-SLAM, a reliability-guided Gaussian–TSDF framework with a sequence-precalibrated prediction stage and an incremental mapping backend. Deterministic reliability gates select admissible observations, while one frozen scale jointly transports predicted depth, camera translation, and length-valued mapping parameters into an internally consistent canonical coordinate. A confidence-weighted colorized TSDF stores the persistent surface, base color, and visibility support. A capacity-bounded Gaussian layer is restricted to reliable residual regions and primarily restores appearance detail over the fused surface. On Replica, Mono3DGS-SLAM achieves 0.249 cm mean ATE and is most accurate in five of eight scenes. With identical poses, depths, and views, the hybrid output reaches 30.678 dB PSNR, exceeding either component alone. ScanNet and TUM-RGBD results further assess transfer to real monocular sequences.
Keywords:
monocular SLAM
; 3D Gaussian Splatting
; TSDF
; bundle adjustment
; canonical dimensional transport
; dense reconstruction
1. Introduction
Reconstructing a geometrically accurate and visually realistic 3D scene from images is a fundamental capability for robot navigation, augmented reality, digital twins, and embodied intelligence. Classical dense SLAM obtains stable surfaces from metric RGB-D measurements through volumetric fusion and global optimization [1,2,3], while neural implicit systems improve scene completion and appearance modeling at the cost of substantial online optimization [4,5,6]. Despite this progress, achieving high-fidelity and geometrically stable reconstruction from a single monocular camera under a bounded mapping budget remains an open challenge.
3D Gaussian Splatting (3DGS) has recently emerged as an efficient explicit representation for novel-view synthesis. Its anisotropic primitives, differentiable rasterization, and visibility-aware rendering provide high-quality appearance with real-time display rates [7]. Subsequent methods improve geometric regularity, anti-aliasing, surface alignment, and scalable scene representation [8,9,10,11]. Gaussian-SLAM, SplaTAM, GS-SLAM, RTG-SLAM, and GS-ICP further incorporate Gaussian primitives into camera tracking and dense mapping [12,13,14,15,16]. However, most of these systems rely on calibrated multi-view input or metric depth, and their geometric stability does not directly transfer to monocular reconstruction.
Monocular 3DGS methods remove the depth sensor by combining learned depth, camera estimation, geometric priors, or photometric optimization. Gaussian Splatting SLAM establishes a unified monocular Gaussian mapping pipeline, while Photo-SLAM, GlORIE-SLAM, Splat-SLAM, and HI-SLAM2 improve tracking, global consistency, and map correction [17,18,19,20,21]. Nevertheless, monocular predictions may remain defined only up to an arbitrary global scene scale, and their depth confidence varies across pixels and time. Residual frame-dependent depth error, accumulated pose drift, occlusion, and appearance changes produce ambiguous photometric residuals, which can be incorrectly absorbed as Gaussian position, scale, or opacity updates. The resulting maps often contain floaters, duplicated surfaces, oversized splats, and incomplete geometry.
Incremental monocular Gaussian mapping introduces an additional systems challenge. Recent approaches employ recurrent tracking, dense factor graphs, keyframe optimization, map deformation, or hierarchical Gaussian maps to extend reconstruction to longer sequences [20,21,22,23,24]. Yet jointly updating poses, depths, and appearance under a finite mapping budget remains poorly conditioned: an incomplete map can become an erroneous photometric attractor, whereas Gaussian insertion and densification, when left unconstrained, can increase computation and memory with sequence length. Hybrid Gaussian–SDF systems improve geometric support [25,26,27], but typically assume metric RGB-D input; directly applying their voxel sizes, spatial gates, Gaussian supports, and learning rates to arbitrary-scale monocular predictions changes the effective geometric meaning of every length-valued operator. For prediction-driven monocular mapping, jointly maintaining observation reliability, dimensional consistency, and bounded hybrid-map growth therefore remains insufficiently studied.
To address these limitations, we propose Mono3DGS-SLAM, a reliability-guided monocular Gaussian–TSDF SLAM system with a sequence-precalibrated prediction stage and an incremental mapping backend. Figure 1 illustrates its central design: a monocular RGB sequence provides camera motion and dense depth, pose-conditioned depth is integrated into a persistent TSDF, and a bounded Gaussian layer refines the rendering in reliable residual regions. Before mapping, covisibility-guided keyframe allocation operates within a fixed budget and supplies selected states to the existing trajectory optimizer. We formulate observation admissibility, dimensional consistency, representation allocation, and persistent capacity as a coupled monocular mapping problem. Our main contributions are:
- A reliability-gated observation formulation and canonical dimensional transport place predicted pose, depth, confidence, intrinsics, and all implemented length-valued mapping operators in one internally consistent, non-metric coordinate without using benchmark geometry.
- An asymmetric Gaussian–TSDF representation assigns persistent surface, base color, and visibility to the TSDF, while a capacity-bounded Gaussian layer is restricted to reliable residual regions and primarily restores appearance detail.
- A controlled evaluation on Replica, ScanNet, and TUM-RGBD combines complete-sequence tracking comparisons with paired trajectory-stage, room0 dimensional-scaling, and eight-scene representation studies under stated protocol limitations.
2. Related Work
2.1. Dense SLAM and Learned Monocular Geometry
Classical dense SLAM systems establish reliable geometry through volumetric fusion, sparse voxel hashing, loop closure, and global surface reintegration [1,2,3]. Neural implicit SLAM improves scene completeness and appearance modeling through learned fields or neural point representations, but still incurs substantial online optimization cost and requires explicit mechanisms for long-term consistency [4,5,6,28]. Learning-based visual odometry and feed-forward reconstruction increasingly estimate camera motion and dense geometry directly from monocular images. DROID-SLAM and DPVO formulate learned recurrent or patch-based geometric optimization [29,30], while DUSt3R, MASt3R, and MASt3R-SLAM reconstruct dense geometry in pointmap space [31,32,33]. For long image streams, Spann3R introduces external spatial memory, whereas VGGT jointly predicts cameras, depth, pointmaps, and tracks [34,35]. VGGT-Long extends feed-forward reconstruction through chunk alignment and loop discovery, and LingBot-Map adopts anchor context, a local pose-reference window, and trajectory memory for causal streaming reconstruction [36,37]. Our work applies input-only scale calibration and geometrically verified pose refinement before dense Gaussian–SDF fusion.
2.2. Gaussian SLAM and Geometry-Aware Reconstruction
3D Gaussian Splatting provides an explicit radiance-field representation with differentiable optimization and real-time rendering [7]. Gaussian-SLAM, SplaTAM, and GS-SLAM use Gaussian primitives for dense tracking, mapping, and novel-view synthesis [12,13,14]. RTG-SLAM improves compact RGB-D reconstruction, while GS-ICP SLAM decouples geometric tracking from Gaussian appearance optimization [15,16]. In the monocular setting, Gaussian Splatting SLAM demonstrates unified Gaussian tracking and mapping [17]. Photo-SLAM combines geometric tracking with Gaussian photometric mapping, whereas Splat-SLAM and HI-SLAM2 introduce monocular priors, global pose optimization, and map deformation [18,20,21]. Complementary work improves geometry through oriented surface primitives, surface alignment, or planar constraints [9,10,11]. Our method instead retains an incrementally fused TSDF as the persistent surface representation and restricts Gaussian creation, visibility, and optimization to geometrically consistent residual regions.
2.3. Hybrid Gaussian–SDF Mapping and Long-Sequence Consistency
Hybrid geometric–radiance representations assign complementary responsibilities to explicit geometry and appearance primitives. GSFusion combines TSDF fusion with Gaussian texture modeling, while GSDF couples SDF-based geometry and Gaussian rendering [25,26]. GPS-SLAM uses a colorized SDF for geometry and base appearance and a sparse Gaussian layer for residual color [27]. These systems assume metrically consistent depth and mainly rely on incremental frontend tracking. Global consistency has been addressed through loop closure, pose-graph optimization, and map deformation in Loopy-SLAM, Splat-SLAM, and HI-SLAM2 [20,21,28]. LoopSplat registers Gaussian submaps directly, while GLC-SLAM performs hierarchical loop closure and submap-level updates [38,39]. GigaSLAM and VINGS-Mono target larger environments [23,24]. Our framework targets pure monocular RGB and studies how prediction-only canonical dimensional transport, shared reliability gating, an existing trajectory optimizer, persistent TSDF geometry, and bounded residual Gaussian appearance can operate in one internally consistent incremental formulation.
3. Method
Given a monocular RGB sequence, Mono3DGS-SLAM reconstructs a refined complete camera trajectory and a hybrid Gaussian–TSDF map. The method is designed around three coupled difficulties of monocular dense mapping: predictions have an arbitrary geometric scale, their reliability varies over pixels and frames, and unconstrained Gaussian parameters can absorb geometric errors as appearance updates. We therefore convert the frontend outputs into reliability-gated, scale-consistent observations, refine the trajectory before persistent fusion, and assign different responsibilities to the two map representations. The TSDF maintains the persistent surface and visibility support, whereas a bounded set of 3D Gaussians is restricted to reliable residual regions and primarily restores detail that is not sufficiently explained by the fused surface.
3.1. Overview
The complete framework is shown in Figure 2. The monocular frontend first predicts camera poses, dense depth, intrinsics, and per-pixel confidence. Before persistent mapping, invalid or low-confidence measurements are rejected, a fixed canonical coordinate is established from the predicted sequence, and covisibility-guided keyframe optimization refines the camera trajectory. This prediction stage removes the global unit mismatch between the monocular outputs and the metric-sensitive mapping operators, while residual depth errors are handled only through deterministic reliability gates during fusion.
The corrected observations are then processed by an incremental mapping backend. Confidence-weighted TSDF fusion constructs the primary persistent surface, accumulates base color, and provides surface-backed visibility. The Gaussian branch is activated only at reliable residual regions where the current surface-backed rendering lacks sufficient support or exhibits large appearance or depth residuals. Hybrid rendering uses Gaussian appearance where available and falls back to the TSDF elsewhere. Local/global view sampling, bounded insertion, scale constraints, and pruning limit both optimization cost and persistent map growth.
3.2. Problem Formulation
For an ordered monocular sequence
the frontend provides
where is a camera-to-world pose, is a dense depth map, is a per-pixel prediction-confidence map, and contains the camera intrinsics. We seek a refined trajectory and a hybrid map
where is a colorized truncated signed-distance field (TSDF) and is a bounded set of 3D Gaussians. The TSDF is the primary persistent surface representation, whereas is a restricted residual layer rather than an independent complete geometry. Benchmark poses and depths are not used for optimization, mapping, gating, or hyperparameter selection.
Monocular predictions have an arbitrary global scale. Direct back-projection therefore gives
only up to an unknown geometric unit. Residual pose error and nonuniform depth error may still place repeated observations of the same surface at different locations. In a Gaussian-only map, these errors can produce floating or oversized primitives; in a TSDF, they lead to the fusion of incompatible signed distances. Our formulation therefore jointly determines the admissible observations, a common dimensional coordinate, the refined trajectory, the persistent TSDF surface, and a finite residual Gaussian state.
3.3. Reliability-Gated Initialization
A persistent map should not be updated by every finite monocular prediction. We use binary gates to determine observation admissibility and continuous weights to control the influence of accepted measurements. A pixel is valid only when its predicted depth is finite and positive and its confidence exceeds . Its confidence weight is
where denotes the sigmoid function. The lowest-confidence valid quantile is additionally removed from the depth priors used by the trajectory optimizer. Strong depth discontinuities are detected from local depth differences and dilated by one pixel to reduce foreground–background mixing.
Frame-level reliability is evaluated from the relative motion between consecutive accepted poses. Let , with translation magnitude and rotation angle . The pose reliability is
A frame with zero pose reliability is excluded from both TSDF fusion and Gaussian insertion. For an accepted corrected-depth pixel, the mapping weight is
where the range-decay term reduces the influence of distant predictions and is the shared binary depth-edge gate. The same admissibility rule governs both map branches so that unreliable evidence cannot enter the geometry and appearance representations through different paths.
3.4. Canonical Dimensional Transport
Monocular scale affects more than the depth map. Scaling depth without changing voxel resolution, truncation distance, Gaussian support, spatial gates, and translation-related update steps changes the effective reconstruction regime. We therefore estimate one fixed canonical scale from the frozen sequence predictions and transport all implemented length-valued quantities together.
Let denote valid depth samples collected over the sequence. We compute the robust statistic
and define
The corrected observation and camera translation are
Rotations and intrinsics remain unchanged. The same factor is applied to TSDF resolution and truncation, view-frustum bounds, depth and spatial thresholds, Gaussian initialization and pruning scales, and translation-related optimization quantities. Dimensionless confidence thresholds and photometric loss weights remain unchanged. This coordinated transport defines an internally consistent geometric unit rather than recovering absolute metric scale, and it does not remove frame-dependent or nonuniform depth deformation.
3.5. Geometry–Appearance Decoupled Mapping
An unconstrained Gaussian map can absorb pose and depth error into primitive centers, scales, and opacities when it is required to explain both geometry and appearance. A TSDF provides a more stable fused surface, but voxel-color averaging and ray casting attenuate high-frequency image detail. We therefore assign asymmetric responsibilities: stores persistent geometry, base color, and visibility support, while is restricted to reliable residual regions and primarily restores appearance detail over the fused surface.
3.5.1. TSDF Geometry and Base Appearance
Admissible, canonically scaled RGB–D-like observations are integrated into a voxel-hashed TSDF. For voxel center , projected pixel , and camera-space voxel depth , the truncated signed-distance observation is
The confidence-weighted update is
where is the bounded fusion weight derived from . Updates are applied only to finite, view-frustum-valid, admissible pixels, and color is accumulated in the same voxel structure. The resulting TSDF acts as the primary persistent surface used for geometric support, visibility, and fallback rendering.
3.5.2. Residual Gaussian Insertion and Rendering
Gaussians are introduced only where reliable image evidence is not adequately explained by the current surface-backed rendering. Let and be the rendered color and depth, the accumulated Gaussian support, and the shared valid-depth, reliability, edge, and TSDF-support gate. A compact candidate rule is
where and are photometric and depth residuals and denotes the local depth-gradient magnitude. A selected center is initialized by
Its footprint is tied to the projected pixel footprint and corrected depth, and only a bounded subsample of candidates is instantiated. TSDF ray casting provides a surface depth for visibility culling; splats that lie sufficiently behind the current fused surface are discarded. Gaussian rendering is used where splat support exists, while TSDF color supplies a fallback in uncovered regions. Thus the Gaussian layer complements the persistent surface instead of reconstructing the entire geometry and base appearance again.
3.6. Backend Optimization
3.6.1. Budgeted Covisibility Refinement
Confidence-gated poses and depth priors initialize the dense trajectory backend. After initial tracking, a geometric covisibility distance is evaluated between neighboring keyframes. Intervals with weak overlap receive additional keyframes, subject to minimum temporal spacing and a fixed global keyframe budget. The resulting set is jointly optimized by the existing robust factor-graph bundle-adjustment solver. This allocation concentrates trajectory refinement where the current motion is weakly constrained without allowing the keyframe graph to grow without bound.
The optimized keyframe poses are aligned to the frontend coordinate convention. Because bundle adjustment covers only keyframes, a motion-only trajectory filler recovers the poses of the remaining frames from neighboring optimized states. The filler does not modify the corrected depth, so the complete trajectory remains in the same canonical coordinate used by the mapping backend.
3.6.2. Bounded Gaussian Refinement
Appearance optimization samples a mixture of recent mapping views and globally distributed keyframes. Recent views recover newly observed detail, whereas historical views reduce appearance forgetting. Reliability determines which residual pixels are eligible for insertion, while fixed per-update and global primitive budgets bound the persistent Gaussian state. Scales are constrained relative to initialization, low-opacity or out-of-range primitives are pruned, and the final global pass fixes the refined trajectory and TSDF while updating only Gaussian appearance and bounded scales.
3.7. Reliability-Aware Objective
For view t, let and denote rendered color and depth. The finite-pixel mask , observation reliability , and an optional edge-aware weight define
The weighted RGB reconstruction loss is
and the structural term is
A weak robust depth term keeps the Gaussian rendering compatible with the fused surface:
The complete Gaussian objective is
where penalizes excessive Gaussian-scale deviation. RGB and SSIM recover high-frequency appearance, the weak depth term maintains compatibility with the persistent surface, and the scale term suppresses oversized splats and floaters. Reliability weighting determines which observations can generate gradients, whereas candidate selection, insertion budgets, and pruning determine which residual state is retained by the map.
4. Experiments
4.1. Experimental Setup
4.1.1. Implementation Details
We implement the reconstruction back end in C++ using LibTorch and custom CUDA kernels for Gaussian rasterization, differentiation, and sparse TSDF fusion. All local experiments are conducted on the same workstation with a single NVIDIA RTX 5090 GPU. For Replica, input images are resized to pixels. The BA prior removes the lowest 15% of finite valid prediction-confidence values independently in each frame. Mapping instead uses a fixed confidence threshold of 1.5 with sigmoid temperature 0.25 and a -dilated depth-edge gate. Covisibility allocation uses threshold 0.05, minimum spacing four frames, and at most 60 additions.
The evaluated scale procedure is a complete-sequence prediction pre-pass: it samples every 20th predicted depth frame and every eighth pixel, maps the valid median to , clips the scale to , and freezes it before mapping. The mapping back end subsequently processes every second input frame, performs 60 local iterations every ten selected frames using four recent views and up to 12 random stored keyframes, and applies two final refinement passes. Before dimensional transport, the TSDF voxel and truncation distances are 0.0075 and 0.030, respectively. Candidate pixels are randomly subsampled at 35%; each update inserts at most 10k Gaussians and the map is capped at 1.2M. The local opacity threshold is 0.003, and the transported maximum scale is . Runtime denotes mapping-backend throughput after the prediction and trajectory-refinement pre-pass. Unless otherwise specified, the same configuration is used throughout a dataset and is not selected using benchmark depth or poses.
4.1.2. Datasets
We evaluate on three public indoor benchmarks: Replica [40], ScanNet v2 [41], and TUM-RGBD [42]. Replica provides photorealistic synthetic sequences with complete geometry and camera-pose annotations. We use all eight standard sequences, each containing 2000 ordered monocular RGB frames, as our principal benchmark for trajectory and appearance evaluation. ScanNet v2 contains real-world handheld RGB-D captures with motion blur, depth noise, and incomplete observations. Following a fixed diagnostic protocol, we evaluate scenes 0000, 0059, and 0106, which represent a long trajectory, challenging camera motion, and a comparatively stable sequence, respectively. These three scenes do not constitute the full ScanNet benchmark. For TUM-RGBD, we use the widely adopted fr1/desk sequence; timestamp association with the official 32 Hz reader yields 592 evaluated frames from 613 source RGB images. Our method receives only the monocular RGB stream in every experiment, and benchmark depth and poses are reserved exclusively for evaluation.
4.1.3. Baselines
We compare against representative monocular Gaussian SLAM systems, HI-SLAM2 and DROID-Splat-Mono, as well as RGB-D Gaussian mapping systems SplaTAM, GS-ICP SLAM, GSFusion, RTG-SLAM, and GauS-SLAM. Replica and TUM results are obtained from local runs whenever executable implementations are available; ScanNet additionally includes published values on the same three scenes, whose provenance is reported in the corresponding table. All local systems are evaluated on the same workstation and use their released configurations, with dataset-specific input paths and camera parameters only.
We report translational ATE RMSE in centimetres, using one evaluation-only alignment for monocular trajectories and alignment for metric RGB-D trajectories. Rendering quality is measured by PSNR, SSIM, and LPIPS. Because released systems differ in view selection and valid-pixel masks, appearance scores obtained under native protocols are treated as diagnostic rather than as a strictly pixel-aligned ranking. Runtime denotes measured mapping-backend throughput and excludes monocular prediction and front-end bundle adjustment unless stated otherwise. TUM results from HI-SLAM2 use its native 612-match evaluator and are marked with a dagger. Column optima are highlighted only when the evaluation protocols are sufficiently comparable.
4.2. Overall Comparison on Three Datasets
4.2.1. Replica
Table 1 provides the controlled tracking comparison: all three RGB-only systems are rerun for 2000 frames per scene and evaluated with the same alignment. Mono3DGS-SLAM obtains a mean ATE of 0.249 cm, compared with 0.258 cm for HI-SLAM2 and 0.286 cm for DROID-Splat-Mono, corresponding to reductions of 3.5% and 12.9%, respectively. The scene-wise profile in Figure 3 shows lower error on room1, office0, office1, office3, and office4. HI-SLAM2 remains more accurate on room0, room2, and office2, indicating that the mean improvement mainly comes from the office sequences. The published RGB-D results provide a metric-depth reference and, as expected, reach lower ATE than the monocular systems. The stage analysis in Table 5 evaluates the complete covisibility-based BA/interpolation block followed by all-frame pose recovery. This ordering is important for mapping: reducing cross-frame pose disagreement before fusion prevents corrected depths from being integrated at inconsistent locations.
Table 2 reports appearance under each method’s native view and masking policy. Mono3DGS-SLAM records the highest reported PSNR (39.678 dB) and SSIM (0.982), improving over HI-SLAM2 by 1.090 dB and 0.010, respectively. HI-SLAM2 attains the lowest LPIPS (0.036 versus 0.045), showing a small perceptual advantage despite the lower pixel fidelity. Mono3DGS-SLAM runs at 32.64 back-end FPS, which is approximately the 10.93 FPS of DROID-Splat-Mono. Overall, the Replica results show that the TSDF backbone and residual Gaussian layer improve pixel reconstruction without requiring dense Gaussian coverage of every surface, enabling real-time back-end operation.
4.2.2. ScanNet
Table 3 evaluates three fixed ScanNet scenes. Mono3DGS-SLAM obtains ATE values of 4.65, 7.12, and 5.65 cm, giving a 5.81 cm average. The corresponding reported averages are 6.64 cm for HI-SLAM2 and 7.26 cm for Splat-SLAM, yielding numerical reductions of 12.5% and 20.0%. The improvement is consistent across the three reported scenes, with the largest absolute reduction over HI-SLAM2 occurring on scene0106 (1.15 cm).
Mono3DGS-SLAM also achieves the highest PSNR at 29.66 dB, exceeding HI-SLAM2 and Splat-SLAM by 1.67 and 1.64 dB. Its LPIPS of 0.234 is comparable to HI-SLAM2 (0.240), whereas Splat-SLAM remains better at 0.173. The SSIM result (0.817) is lower than both monocular baselines, indicating that the PSNR gain is accompanied by weaker structural similarity. This metric pattern is consistent with the proposed division of representation: confidence-weighted TSDF fusion favors continuous low-frequency surfaces, while the bounded Gaussian branch restores residual appearance only where reliable support is available. It improves pixel fidelity but does not recover all fine structures in the more irregular ScanNet geometry.
4.2.3. TUM-RGBD
Table 4 reports the complete fr1/desk experiment. Mono3DGS-SLAM processes all 592 associated frames and obtains 1.651 cm ATE, 25.548 dB PSNR, 0.7552 SSIM, and 0.2040 LPIPS at 37.31 back-end FPS. Its ATE is 10.6% lower than HI-SLAM2 and 8.7% lower than GauS-SLAM. Mono3DGS-SLAM also improves PSNR by 5.116 dB over HI-SLAM2 and 2.015 dB over GauS-SLAM, while reducing LPIPS to 0.2040. GauS-SLAM retains the highest SSIM (0.8954), and GS-ICP provides the highest back-end throughput (111.82 FPS). Thus, our method gives the strongest reported ATE, PSNR, and LPIPS on this sequence, while remaining below the best structural-similarity and speed results. RTG-SLAM does not produce a valid map, and the GPS-SLAM run yields 125.038 cm ATE without a final rendering. The complete 592-frame reconstruction shows that the canonical depth correction and reliability gates remain operational on real monocular input, while the SSIM gap indicates that depth noise and fine structural detail remain less well handled than the dominant appearance.
4.3. Qualitative Comparison
Figure 4 compares reference images, archived HI-SLAM2 outputs, and our hybrid renderings on three Replica scenes. Both methods reproduce the dominant colors and room layout. HI-SLAM2 preserves fine texture in several local regions, whereas our TSDF fallback provides more continuous support near object and wall boundaries when Gaussian coverage is insufficient. Residual Gaussians recover much of the local appearance without defining the complete geometry. The remaining errors are concentrated around thin structures and high-contrast edges, where uncertain depth is downweighted and therefore provides limited support to both branches.
Figure 5 compares RTG-SLAM, GS-ICP, and GSFusion [15,16,25] with our RGB-only map on room0. All four methods recover the enclosing walls and principal furniture. The RGB-D reconstructions exhibit different levels of sparse support and isolated points, while confidence-weighted TSDF fusion forms denser planar surfaces and preserves the major object arrangement from RGB input alone. Some duplicated support remains near open boundaries, showing that canonical dimensional transport does not remove local monocular depth inconsistency.
Figure 6 compares the two independently executed monocular systems on three ScanNet scenes. Both recover the dominant visible content. Our hybrid rendering improves coverage on walls, floors, and large furniture through TSDF fallback, while the residual Gaussian layer recovers color and local texture on supported regions. Fine geometry is nevertheless smoothed in several cluttered areas. This behavior agrees with the higher PSNR and lower SSIM in Table 3: low-frequency appearance and surface continuity improve, while structural detail remains challenging.
The geometry comparisons show the same trend. In Figure 7, our TSDF-coloured maps contain more continuous large surfaces than the two monocular Gaussian maps, although room1 and room2 retain duplicated boundary support. Figure 8 further shows improved coverage of the dominant room structure, together with residual holes and floaters in regions with limited observations. These results indicate that explicit fusion stabilizes large-scale geometry, whereas thin objects, occlusion boundaries, and weakly observed regions remain the primary failure cases. This is consistent with the reliability-aware design: low-confidence measurements do not corrupt the persistent map, but aggressive rejection can leave unsupported geometry that later views do not fill.
Image and model provenance.
Every figure has a separate directory under figures/dataset_qualitative. Its sources.csv records the scene, input modality, representation, semantic role, implementation directory, exact absolute artifact path, and generated local panel. README.md states all resizing, view selection, and PCA projection operations. No panel is synthesized, and missing baseline maps are listed in COVERAGE_AUDIT.md.
4.4. Ablation Experiments
Pose refinement and all-frame recovery.
Table 5 evaluates the implemented trajectory stages under a common Replica protocol. Adding the covisibility-based BA and keyframe-interpolation block reduces mean ATE from 6.087 to 1.357 cm (77.7%). Adding the remaining all-frame completion stage reduces the mean to 0.249 cm, an 81.6% reduction relative to the preceding stage and 95.9% relative to the raw input. The reduction is observed on all eight scenes, with the largest absolute gain appearing on office0. The first refinement stage removes most of the initial trajectory error, while all-frame completion provides the final sub-centimetre accuracy by propagating the optimized motion to non-keyframes. This supports the method design of stabilizing the complete trajectory before committing observations to the TSDF and Gaussian map.
Table 5.
Replica front-end stage ablation, ATE RMSE in cm. The best completed stage is bold.
| Stage | room0 | room1 | room2 | office0 | office1 | office2 | office3 | office4 | Avg. |
| Raw data | 4.831 | 6.277 | 8.087 | 9.108 | 3.866 | 5.969 | 5.760 | 4.795 | 6.087 |
| Covisibility-guided BA + interpolation | 1.268 | 1.331 | 1.147 | 1.794 | 1.337 | 1.162 | 1.358 | 1.455 | 1.357 |
| All | 0.242 | 0.218 | 0.213 | 0.190 | 0.248 | 0.277 | 0.312 | 0.294 | 0.249 |
Canonical depth and parameter scaling.
Table 6 gives the available paired room0 comparison for canonical scaling. Relative to fixed spatial parameters, transporting the corrected depth scale consistently through spatial thresholds and coverage settings increases PSNR by 6.164 dB and SSIM by 0.040, and decreases LPIPS by 0.037. This improvement is accompanied by a decrease in measured back-end rate from 12.94 to 12.25 FPS (5.3%); peak GPU memory changes from 21,629 to 21,283 MB (1.6% lower). The large appearance gain with nearly unchanged memory shows that applying one scale to depth, translation, TSDF support, and Gaussian spatial parameters avoids a mismatched reconstruction regime.
TSDF–Gaussian representation.
All rows in Table 7 use the same runs, poses, depths, and evaluation views, providing the most directly paired component comparison. TSDF raycasting increases PSNR from 18.103 to 24.680 dB (+6.577 dB) relative to raw 3DGS, while also improving SSIM from 0.703 to 0.822 and LPIPS from 0.347 to 0.301. Combining the two representations raises PSNR to 30.678 dB and SSIM to 0.899, and lowers LPIPS to 0.193. Relative to TSDF alone, these changes are +5.998 dB, +0.077, and , respectively.
Figure 9 shows that the PSNR ordering is consistent across the eight evaluated scenes. The examples in Figure 10 illustrate the corresponding trade-off: raw Gaussians preserve local texture but can leave coverage artifacts, TSDF provides smoother continuous support, and the hybrid restores part of the appearance detail. The consistent metric improvements show that the two representations contribute complementary information: TSDF supplies stable surface support and visibility, while residual Gaussians recover image detail that is attenuated by raycasting alone. The result also agrees with the reliability-aware objective: RGB and structural losses refine appearance on valid support, whereas the weak depth and scale terms keep the Gaussian layer compatible with the fixed geometric backbone.
4.5. Limitations
The principal limitation is that canonical scale is estimated from a complete-sequence prediction pre-pass. The evaluated pipeline is therefore not strictly causal, its coordinate remains non-metric, and one scale cannot correct spatially or temporally varying depth errors. Deterministic gates suppress weak observations but neither recover rejected geometry nor prevent accepted errors from entering the TSDF, particularly around thin or weakly observed structures and in dynamic scenes. Bounded Gaussian capacity also does not eliminate the substantial memory footprint, while the reported throughput measures the mapping backend rather than the complete pipeline. Finally, native appearance protocols, limited ScanNet and TUM coverage, and coupled-stage ablations restrict strict cross-system and component-level conclusions; broader fixed-protocol, multi-scene end-to-end evaluation is required.
5. Conclusion
This work presented Mono3DGS-SLAM, a reliability-gated formulation for monocular Gaussian–TSDF reconstruction. Its central contribution is to place predicted pose, depth, confidence, intrinsics, and length-valued mapping operators in a shared canonical coordinate, and to separate persistent geometry from bounded residual appearance. This coupling limits the propagation of weak monocular observations and unconstrained Gaussian growth during incremental mapping. On eight Replica scenes, Mono3DGS-SLAM obtains 0.249 cm mean ATE compared with 0.258 cm for HI-SLAM2. Under identical poses, depths, and views, the hybrid representation reaches 30.678 dB PSNR, versus 18.103 dB for raw 3DGS and 24.680 dB for TSDF raycasting, supporting the complementary roles of its two map representations. Selected ScanNet and TUM-RGBD results indicate transfer to real monocular sequences while exposing sensitivity to local depth errors and resource demands. Future work will pursue causal scale estimation, lower end-to-end latency, and broader fixed-protocol evaluation.
Author Contributions
Conceptualization, Y.W. and B.W.; methodology, Y.W. and B.W.; software, Y.W.; validation, Y.W., B.W., J.X. and G.Y.; formal analysis, Y.W. and B.W.; investigation, Y.W.; data curation, Y.W.; writing—original draft preparation, Y.W. and B.W.; writing—review and editing, Y.W., B.W., J.X., G.Y. and H.W.; visualization, Y.W.; resources, J.X.; supervision, B.W. and H.W.; project administration, Y.W. and B.W.; funding acquisition, H.W., B.W. and J.X. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the Shandong Provincial Natural Science Foundation, grant number ZR2024QE098, and the Liaoning Provincial Science and Technology Joint Program (Natural Science Foundation–Doctoral Research Start-up Project), grant number 2025-BSLH-333.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The evaluation archive contains saved trajectories, configurations, Gaussian models, TSDF meshes, colored point clouds, rendering samples, and machine-readable metric tables. Data will be made available on request.
Conflicts of Interest
The authors declare no conflict of interest.
References
- Newcombe, R.A.; Izadi, S.; Hilliges, O.; Molyneaux, D.; Kim, D.; Davison, A.J.; Kohli, P.; Shotton, J.; Hodges, S.; Fitzgibbon, A. KinectFusion: Real-Time Dense Surface Mapping and Tracking. Proceedings of the 10th IEEE International Symposium on Mixed and Augmented Reality (ISMAR), 2011, pp. 127–136. [CrossRef]
- Nießner, M.; Zollhöfer, M.; Izadi, S.; Stamminger, M. Real-Time 3D Reconstruction at Scale Using Voxel Hashing. ACM Transactions on Graphics 2013, 32, 169:1–169:11. [CrossRef]
- Dai, A.; Nießner, M.; Zollhöfer, M.; Izadi, S.; Theobalt, C. BundleFusion: Real-Time Globally Consistent 3D Reconstruction Using On-the-Fly Surface Reintegration. ACM Transactions on Graphics 2017, 36, 76:1–76:18. [CrossRef]
- Zhu, Z.; Peng, S.; Larsson, V.; Xu, W.; Bao, H.; Cui, Z.; Oswald, M.R.; Pollefeys, M. NICE-SLAM: Neural Implicit Scalable Encoding for SLAM. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 12786–12796. [CrossRef]
- Wang, H.; Wang, J.; Agapito, L. Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 13293–13302. [CrossRef]
- Sandström, E.; Li, Y.; Van Gool, L.; Oswald, M.R. Point-SLAM: Dense Neural Point Cloud-Based SLAM. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 18433–18444. [CrossRef]
- Kerbl, B.; Kopanas, G.; Leimkühler, T.; Drettakis, G. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics 2023, 42, 139:1–139:14. [CrossRef]
- Yu, Z.; Chen, A.; Huang, B.; Sattler, T.; Geiger, A. Mip-Splatting: Alias-Free 3D Gaussian Splatting. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, [arXiv:cs.CV/2311.16493].
- Huang, B.; Yu, Z.; Chen, A.; Geiger, A.; Gao, S. 2D Gaussian Splatting for Geometrically Accurate Radiance Fields. ACM SIGGRAPH 2024 Conference Papers, 2024, [arXiv:cs.CV/2403.17888].
- Guédon, A.; Lepetit, V. SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction and High-Quality Mesh Rendering. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5354–5363, [arXiv:cs.CV/2311.12775].
- Chen, D.; Li, H.; Ye, W.; Wang, Y.; Xie, W.; Zhai, S.; Wang, N.; Liu, H.; Bao, H.; Zhang, G. PGSR: Planar-Based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction. arXiv preprint arXiv:2406.06521 2024, [arXiv:cs.CV/2406.06521].
- Yugay, V.; Li, Y.; Gevers, T.; Oswald, M.R. Gaussian-SLAM: Photo-Realistic Dense SLAM with Gaussian Splatting. arXiv preprint arXiv:2312.10070 2023, [arXiv:cs.CV/2312.10070].
- Keetha, N.; Karhade, J.; Jatavallabhula, K.M.; Yang, G.; Scherer, S.; Ramanan, D.; Luiten, J. SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
- Yan, C.; Qu, D.; Xu, D.; Zhao, B.; Wang, Z.; Wang, D.; Li, X. GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 19595–19604. [CrossRef]
- Peng, Z.; Shao, T.; Liu, Y.; Zhou, J.; Yang, Y.; Wang, J.; Zhou, K. RTG-SLAM: Real-Time 3D Reconstruction at Scale Using Gaussian Splatting. ACM SIGGRAPH 2024 Conference Papers, 2024. [CrossRef]
- Ha, S.; Yeon, J.; Yu, H. RGBD GS-ICP SLAM. Proceedings of the European Conference on Computer Vision (ECCV), 2024, pp. 180–197. [CrossRef]
- Matsuki, H.; Murai, R.; Kelly, P.H.J.; Davison, A.J. Gaussian Splatting SLAM. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 18039–18048.
- Huang, H.; Li, L.; Cheng, H.; Yeung, S.K. Photo-SLAM: Real-Time Simultaneous Localization and Photorealistic Mapping for Monocular, Stereo, and RGB-D Cameras. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
- Zhang, G.; Sandström, E.; Zhang, Y.; Patel, M.; Van Gool, L.; Oswald, M.R. GlORIE-SLAM: Globally Optimized RGB-Only Implicit Encoding Point Cloud SLAM. arXiv preprint arXiv:2403.19549 2024, [arXiv:cs.CV/2403.19549].
- Sandström, E.; Tateno, K.; Oechsle, M.; Niemeyer, M.; Van Gool, L.; Oswald, M.R.; Tombari, F. Splat-SLAM: Globally Optimized RGB-Only SLAM with 3D Gaussians. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025, [arXiv:cs.CV/2405.16544].
- Zhang, W.; Cheng, Q.; Skuddis, D.; Zeller, N.; Cremers, D.; Haala, N. HI-SLAM2: Geometry-Aware Gaussian SLAM for Fast Monocular Scene Reconstruction. IEEE Transactions on Robotics 2025, 41, 6478–6493, [arXiv:cs.CV/2411.17982].
- Homeyer, C.; Begiristain, L.; Schnörr, C. DROID-Splat: Combining End-to-End SLAM with 3D Gaussian Splatting. Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2025.
- Deng, K.; Yang, J.; Wang, S.; Xie, J. GigaSLAM: Large-Scale Monocular SLAM with Hierarchical Gaussian Splats. arXiv preprint arXiv:2503.08071 2025, [arXiv:cs.CV/2503.08071].
- Wu, K.; Zhang, Z.; Tie, M.; Ai, Z.; Gan, Z.; Ding, W. VINGS-Mono: Visual-Inertial Gaussian Splatting Monocular SLAM in Large Scenes. arXiv preprint arXiv:2501.08286 2025, [arXiv:cs.RO/2501.08286].
- Wei, J.; Leutenegger, S. GSFusion: Online RGB-D Mapping Where Gaussian Splatting Meets TSDF Fusion. IEEE Robotics and Automation Letters 2024, 9, 11865–11872. [CrossRef]
- Yu, M.; Lu, T.; Xu, L.; Jiang, L.; Xiangli, Y.; Dai, B. GSDF: 3DGS Meets SDF for Improved Rendering and Reconstruction. Advances in Neural Information Processing Systems 2024, 37, [arXiv:cs.CV/2403.16964].
- Peng, Z.; Zhou, K.; Shao, T. Gaussian-plus-SDF SLAM: High-Fidelity 3D Reconstruction at 150+ FPS. Computational Visual Media 2025, 11, 1195–1208. [CrossRef]
- Liso, L.; Sandström, E.; Yugay, V.; Van Gool, L.; Oswald, M.R. Loopy-SLAM: Dense Neural SLAM with Loop Closures. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 20363–20373.
- Teed, Z.; Deng, J. DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. Advances in Neural Information Processing Systems, 2021, Vol. 34, pp. 16558–16569.
- Teed, Z.; Lipson, L.; Deng, J. Deep Patch Visual Odometry. Advances in Neural Information Processing Systems, 2022, Vol. 35, pp. 39035–39047.
- Wang, S.; Leroy, V.; Cabon, Y.; Chidlovskii, B.; Revaud, J. DUSt3R: Geometric 3D Vision Made Easy. arXiv preprint arXiv:2312.14132 2023, [arXiv:cs.CV/2312.14132].
- Leroy, V.; Cabon, Y.; Revaud, J. Grounding Image Matching in 3D with MASt3R. Proceedings of the European Conference on Computer Vision, 2024, [arXiv:cs.CV/2406.09756].
- Murai, R.; Dexheimer, E.; Davison, A.J. MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors. arXiv preprint arXiv:2412.12392 2024, [arXiv:cs.CV/2412.12392].
- Wang, H.; Agapito, L. 3D Reconstruction with Spatial Memory. arXiv preprint arXiv:2408.16061 2024, [arXiv:cs.CV/2408.16061].
- Wang, J.; Chen, M.; Karaev, N.; Vedaldi, A.; Rupprecht, C.; Novotny, D. VGGT: Visual Geometry Grounded Transformer. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, [arXiv:cs.CV/2503.11651].
- Deng, K.; Ti, Z.; Xu, J.; Yang, J.; Xie, J. VGGT-Long: Chunk It, Loop It, Align It—Pushing VGGT’s Limits on Kilometer-Scale Long RGB Sequences. arXiv preprint arXiv:2507.16443 2025, [arXiv:cs.CV/2507.16443].
- Chen, L.Z.; Gao, J.; Chen, Y.; Cheng, K.L.; Sun, Y.; Hu, L.; Xue, N.; Zhu, X.; Shen, Y.; Yao, Y.; et al. Geometric Context Transformer for Streaming 3D Reconstruction. arXiv preprint arXiv:2604.14141 2026, [arXiv:cs.CV/2604.14141].
- Zhu, L.; Li, Y.; Sandström, E.; Huang, S.; Schindler, K.; Armeni, I. LoopSplat: Loop Closure by Registering 3D Gaussian Splats. arXiv preprint arXiv:2408.10154 2024, [arXiv:cs.CV/2408.10154].
- Xu, Z.; Li, Q.; Chen, C.; Liu, X.; Niu, J. GLC-SLAM: Gaussian Splatting SLAM with Efficient Loop Closure. arXiv preprint arXiv:2409.10982 2024, [arXiv:cs.RO/2409.10982].
- Straub, J.; Whelan, T.; Ma, L.; Chen, Y.; Wijmans, E.; Green, S.; Engel, J.J.; Mur-Artal, R.; Ren, C.; Verma, S.; et al. The Replica Dataset: A Digital Replica of Indoor Spaces. arXiv preprint arXiv:1906.05797 2019.
- Dai, A.; Chang, A.X.; Savva, M.; Halber, M.; Funkhouser, T.; Nießner, M. ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
- Sturm, J.; Engelhard, N.; Endres, F.; Burgard, W.; Cremers, D. A Benchmark for the Evaluation of RGB-D SLAM Systems. Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, 2012.
- Su, Y.; Chen, L.; Zhang, K.; Zhao, Z.; Hou, C.; Yu, Z. GauS-SLAM: Dense RGB-D SLAM with Gaussian Surfels. arXiv preprint arXiv:2505.01934 2025, [arXiv:cs.CV/2505.01934].
Figure 1.
System concept and representative reconstruction. From left to right, a monocular RGB sequence is processed to obtain camera poses and dense depth. The pose-conditioned depth observations are fused into a dense TSDF, while the image observations support a complementary 3D Gaussian appearance representation. Hybrid rendering uses Gaussian-supported appearance where available and TSDF color as fallback elsewhere. The RGB frames, depth visualization, and reconstructed scene are actual results from Replica room0.
Figure 1.
System concept and representative reconstruction. From left to right, a monocular RGB sequence is processed to obtain camera poses and dense depth. The pose-conditioned depth observations are fused into a dense TSDF, while the image observations support a complementary 3D Gaussian appearance representation. Hybrid rendering uses Gaussian-supported appearance where available and TSDF color as fallback elsewhere. The RGB frames, depth visualization, and reconstructed scene are actual results from Replica room0.

Figure 2.
Method formulation of Mono3DGS-SLAM. A sequence prediction pre-pass establishes observation admissibility, refined poses, and a fixed canonical coordinate; the mapping backend then updates the map incrementally. Admissibility controls evidence, dimensional transport controls geometric units, the TSDF/Gaussian decomposition assigns persistent-surface and residual-region responsibilities, and fixed budgets bound Gaussian insertion, refinement, and persistent state. Solid arrows show data flow; dashed arrows show geometric or optimization constraints.
Figure 2.
Method formulation of Mono3DGS-SLAM. A sequence prediction pre-pass establishes observation admissibility, refined poses, and a fixed canonical coordinate; the mapping backend then updates the map incrementally. Admissibility controls evidence, dimensional transport controls geometric units, the TSDF/Gaussian decomposition assigns persistent-surface and residual-region responsibilities, and fixed budgets bound Gaussian insertion, refinement, and persistent state. Solid arrows show data flow; dashed arrows show geometric or optimization constraints.

Figure 3.
Per-scene Replica ATE for three locally rerun monocular systems. All rows use complete 2000-frame trajectories and the same evaluation alignment.
Figure 3.
Per-scene Replica ATE for three locally rerun monocular systems. All rows use complete 2000-frame trajectories and the same evaluation alignment.

Figure 4.
Replica appearance comparison on room0, room1, and room2 (top to bottom). Columns show the reference image, the nearest saved HI-SLAM2 output, and the Gaussian–TSDF hybrid rendering of Mono3DGS-SLAM (left to right). Reference and Mono3DGS-SLAM use frame 666; HI-SLAM2 uses its nearest archived views (655, 655, and 660, respectively). Red boxes highlight representative local differences in texture and boundary reconstruction.
Figure 4.
Replica appearance comparison on room0, room1, and room2 (top to bottom). Columns show the reference image, the nearest saved HI-SLAM2 output, and the Gaussian–TSDF hybrid rendering of Mono3DGS-SLAM (left to right). Reference and Mono3DGS-SLAM use frame 666; HI-SLAM2 uses its nearest archived views (655, 655, and 660, respectively). Red boxes highlight representative local differences in texture and boundary reconstruction.

Figure 5.
Replica room0 geometry comparison between three RGB-D baselines and our RGB-only method. RTG-SLAM, GS-ICP, and GSFusion visualize their native final Gaussian maps, whereas Mono3DGS-SLAM shows its TSDF-coloured point cloud. All panels use the same canonical orbit view (elevation , azimuth ) after independent robust centring and PCA alignment. The comparison therefore emphasizes structural completeness and artifacts rather than metric scale or a shared world-space camera pose.
Figure 5.
Replica room0 geometry comparison between three RGB-D baselines and our RGB-only method. RTG-SLAM, GS-ICP, and GSFusion visualize their native final Gaussian maps, whereas Mono3DGS-SLAM shows its TSDF-coloured point cloud. All panels use the same canonical orbit view (elevation , azimuth ) after independent robust centring and PCA alignment. The comparison therefore emphasizes structural completeness and artifacts rather than metric scale or a shared world-space camera pose.

Figure 6.
ScanNet v2 cross-method rendering comparison on scene0059_00, scene0000_00, and scene0106_00 (top to bottom), using frames 1786, 1332, and 1998, respectively. Each row collects the archived reference image, the independently executed HI-SLAM2 optimized Gaussian rendering, and the final Gaussian–TSDF hybrid rendering of Mono3DGS-SLAM.
Figure 6.
ScanNet v2 cross-method rendering comparison on scene0059_00, scene0000_00, and scene0106_00 (top to bottom), using frames 1786, 1332, and 1998, respectively. Each row collects the archived reference image, the independently executed HI-SLAM2 optimized Gaussian rendering, and the final Gaussian–TSDF hybrid rendering of Mono3DGS-SLAM.

Figure 7.
Replica cross-method geometry on room0, room1, and room2. DROID-Splat and HI-SLAM2 are shown from their native final PLY maps, and Mono3DGS-SLAM is shown from its final TSDF-coloured point cloud. Independent robust centring and PCA projection make this a qualitative completeness/artifact comparison rather than a common-view metric evaluation.
Figure 7.
Replica cross-method geometry on room0, room1, and room2. DROID-Splat and HI-SLAM2 are shown from their native final PLY maps, and Mono3DGS-SLAM is shown from its final TSDF-coloured point cloud. Independent robust centring and PCA projection make this a qualitative completeness/artifact comparison rather than a common-view metric evaluation.

Figure 8.
ScanNet v2 cross-method 3D/TSDF comparison on scene0000_00, scene0059_00, and scene0106_00. Columns show the ScanNet clean mesh, HI-SLAM2 geometry fused from its optimized rendered depths, and the native TSDF-coloured map of Mono3DGS-SLAM. Each panel is deterministically and independently centred and PCA-projected, so the comparison emphasizes structural coverage rather than absolute scale.
Figure 8.
ScanNet v2 cross-method 3D/TSDF comparison on scene0000_00, scene0059_00, and scene0106_00. Columns show the ScanNet clean mesh, HI-SLAM2 geometry fused from its optimized rendered depths, and the native TSDF-coloured map of Mono3DGS-SLAM. Each panel is deterministically and independently centred and PCA-projected, so the comparison emphasizes structural coverage rather than absolute scale.

Figure 9.
Quantitative Replica ablation. Panel (a) shows trajectory accuracy at both implemented refinement stages. Panel (b) shows that the hybrid output exceeds raw 3DGS and TSDF-only PSNR in every scene.
Figure 9.
Quantitative Replica ablation. Panel (a) shows trajectory accuracy at both implemented refinement stages. Panel (b) shows that the hybrid output exceeds raw 3DGS and TSDF-only PSNR in every scene.

Figure 10.
Paired qualitative ablation on room0, office2, and office3. Raw Gaussians retain local texture but exhibit coverage and boundary artifacts; TSDF provides smoother surface support; in these examples, the full hybrid restores appearance detail while reducing visible holes and black boundary artifacts.
Figure 10.
Paired qualitative ablation on room0, office2, and office3. Raw Gaussians retain local texture but exhibit coverage and boundary artifacts; TSDF provides smoother surface support; in these examples, the full hybrid restores appearance detail while reducing visible holes and black boundary artifacts.

Table 1.
Replica camera tracking comparison (ATE RMSE in cm). RGB rows are local 2000-frame runs with evaluation-only alignment. RGB-D rows are values reported in the respective papers under method-native metric protocols and are not strictly comparable to the RGB block. Best complete results within each input block are bold. †GPS-SLAM reports only room0 and the eight-scene mean; per-scene values for the other sequences are not published.
Table 1.
Replica camera tracking comparison (ATE RMSE in cm). RGB rows are local 2000-frame runs with evaluation-only alignment. RGB-D rows are values reported in the respective papers under method-native metric protocols and are not strictly comparable to the RGB block. Best complete results within each input block are bold. †GPS-SLAM reports only room0 and the eight-scene mean; per-scene values for the other sequences are not published.
| Input | Metdod | room0 | room1 | room2 | office0 | office1 | office2 | office3 | office4 | Avg. |
| RGB-D | SplaTAM [13] | 0.31 | 0.40 | 0.29 | 0.47 | 0.27 | 0.29 | 0.32 | 0.55 | 0.36 |
| RGB-D | GS-ICP SLAM [16] | 0.15 | 0.16 | 0.11 | 0.18 | 0.12 | 0.17 | 0.16 | 0.21 | 0.16 |
| RGB-D | RTG-SLAM [15] | 0.20 | 0.18 | 0.13 | 0.22 | 0.12 | 0.22 | 0.20 | 0.19 | 0.18 |
| RGB-D | GauS-SLAM [43] | 0.06 | 0.08 | 0.08 | 0.06 | 0.03 | 0.09 | 0.05 | 0.05 | 0.06 |
| RGB-D | GPS-SLAM† [27] | 0.16 | – | – | – | – | – | – | – | 0.17 |
| RGB | HI-SLAM2 [21] | 0.228 | 0.220 | 0.185 | 0.223 | 0.286 | 0.259 | 0.324 | 0.341 | 0.258 |
| RGB | DROID-Splat-Mono [22] | 0.280 | 0.250 | 0.208 | 0.200 | 0.296 | 0.296 | 0.355 | 0.398 | 0.286 |
| RGB | Mono3DGS-SLAM | 0.242 | 0.218 | 0.213 | 0.190 | 0.248 | 0.277 | 0.312 | 0.294 | 0.249 |
Table 2.
Replica eight-scene native-protocol appearance and throughput. Cross-method rendering masks and view policies differ; values are therefore descriptive. Best reported numeric values in each column are bold.
Table 2.
Replica eight-scene native-protocol appearance and throughput. Cross-method rendering masks and view policies differ; values are therefore descriptive. Best reported numeric values in each column are bold.
| Metdod | Input | PSNR↑ | SSIM↑ | LPIPS↓ | Backend FPS↑ |
| HI-SLAM2 [21] | RGB | 38.588 | 0.972 | 0.036 | – |
| DROID-Splat-Mono [22] | RGB | 33.039 | 0.928 | 0.117 | 10.93 |
| Mono3DGS-SLAM | RGB | 39.678 | 0.982 | 0.045 | 32.64 |
Table 3.
ScanNet selected-scene comparison. ATE is in cm. Baselines are literature values on scenes 0000, 0059, and 0106; ours is a local rerun. Appearance columns are native-protocol three-scene averages. The mixed provenance precludes a strict cross-row ranking; bold denotes only the best reported numeric value in each column.
Table 3.
ScanNet selected-scene comparison. ATE is in cm. Baselines are literature values on scenes 0000, 0059, and 0106; ours is a local rerun. Appearance columns are native-protocol three-scene averages. The mixed provenance precludes a strict cross-row ranking; bold denotes only the best reported numeric value in each column.
| Input | Metdod | 0000 | 0059 | 0106 | ATE Avg.↓ | PSNR↑ | SSIM↑ | LPIPS↓ |
| RGB | HI-SLAM2 [21] | 5.82 | 7.30 | 6.80 | 6.64 | 27.99 | 0.873 | 0.240 |
| RGB | Splat-SLAM [20] | 5.57 | 9.11 | 7.09 | 7.26 | 28.02 | 0.853 | 0.173 |
| RGB-D | SplaTAM [13] | 12.80 | 10.10 | 17.70 | 13.53 | 19.82 | 0.770 | 0.373 |
| RGB-D | Gaussian-SLAM [12] | 21.20 | 12.80 | 13.50 | 15.83 | 27.00 | 0.930 | 0.233 |
| RGB | Mono3DGS-SLAM | 4.65 | 7.12 | 5.65 | 5.81 | 29.66 | 0.817 | 0.234 |
Table 4.
TUM-RGBD fr1/desk local comparison. †HI-SLAM2 uses its native 612-match protocol; the remaining completed rows use the same 592 associated frames. Best values among valid rows are bold. FPS is back-end throughput, not end-to-end throughput. ‡GPS-SLAM reports pipeline FPS, but its invalid final map is excluded from ranking.
Table 4.
TUM-RGBD fr1/desk local comparison. †HI-SLAM2 uses its native 612-match protocol; the remaining completed rows use the same 592 associated frames. Best values among valid rows are bold. FPS is back-end throughput, not end-to-end throughput. ‡GPS-SLAM reports pipeline FPS, but its invalid final map is excluded from ranking.
| Input | Metdod | Frames | Align. | ATE (cm)↓ | PSNR↑ | SSIM↑ | LPIPS↓ | FPS↑ |
| RGB | HI-SLAM2† [21] | 612 | 1.846 | 20.432 | 0.7503 | 0.2367 | – | |
| RGB-D | GS-ICP SLAM [16] | 592 | 2.975 | 17.194 | 0.6965 | 0.3232 | 111.82 | |
| RGB-D | GauS-SLAM [43] | 592 | 1.808 | 23.533 | 0.8954 | 0.2369 | – | |
| RGB-D | RTG-SLAM [15] | – | – | – | – | – | – | |
| RGB-D | GPS-SLAM‡ [27] | 592 | 125.038 | – | – | – | 91.03 | |
| RGB | Mono3DGS-SLAM | 592 | 1.651 | 25.548 | 0.7552 | 0.2040 | 37.31 |
Table 6.
Room0 paired comparison of canonical dimensional scaling.
| Method | PSNR↑ | SSIM↑ | LPIPS↓ | FPS↑ | GPU MB↓ |
| Fixed spatial parameters | 20.554 | 0.769 | 0.257 | 12.94 | 21629 |
| Unified dimensional scaling | 26.718 | 0.809 | 0.220 | 12.25 | 21283 |
Table 7.
Eight-scene paired representation ablation.
| Output | PSNR↑ | SSIM↑ | LPIPS↓ |
| Raw 3DGS | 18.103 | 0.703 | 0.347 |
| TSDF raycast | 24.680 | 0.822 | 0.301 |
| Full Mono3DGS-SLAM Hybrid | 30.678 | 0.899 | 0.193 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.