SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Shing Ho J. Lin*
Wenzhao Zheng*, †
Dong Zhuo*
Yuqi Wu
Jie Zhou
Jiwen Lu
Intelligent Vision Group, Tsinghua University
*Equal Contributions   Project Leader  

Visualization of Our Method

(Point Clouds are Downsampled 4 Times for Efficiency)

SM4RT Video
Loading video...
SM4RT Synchronized 3D Point Cloud
🎬swing/frame_*.ply📁 0 frames
0 / 0
Speed1.0x
Point Size0.02

Method Comparison

(Point Clouds are Downsampled 4 Times for Efficiency)

Scene
Modality
Baseline
SM4RT Ours Video
Loading SM4RT...
swing Preparing...
4RC Video
Loading 4RC...
swing_4rc Preparing...
1 / 24

Overview: From Geometry-with-Motion to Geometry-of-Motion

SM4RT teaser

Figure 1: SM4RT decomposes scene motion into structured latent bases.

Recent advances in Geometry Foundation Models (GFMs) have demonstrated remarkable capability in reconstructing 3D scene geometry from monocular video. With static geometry reconstruction now approaching practical fidelity, the next frontier is 4D dynamic understanding which jointly infers scene structure and the underlying motion.

We propose SM4RT, a Structured Motion 4D Reconstruction Transformer framework. Rather than predicting independent point-wise displacements, SM4RT introduces Structure-of-Motion (SoM): scene motion is composed of a compact set of N latent motion bases, each represented as a temporal sequence of 6D se(3) transforms.

Motivation1: Humans perceive motion via sparse kinematic cues—not at the pixel level, but at the level of objects and parts. Real-world objects usually obey rigid-body kinematics: points move collectively, not in isolation.
Motivation2: Motion exhibits geometric structure as fiber bundle, formalizing scene dynamics as a principal bundle whose local sections correspond to valid rigid-body motions, ensuring that point-wise displacements respect the underlying topological and kinematic constraints.
comparison

Figure 2: Comparison of motion representations. The proposed representation preserves the structure of object motion, providing a more structured description than standard per-pixel displacement and encouraging intra-object homogeneity.

Overall Framework of SM4RT

SM4RT Architecture

Figure 3: Architecture of SM4RT.

Given a monocular input video, SM4RT extracts latent geometry tokens with a DINOv2 backbone, then decouples them into scene geometry and motion geometry tokens. The scene decoder predicts camera parameters and depth, while the motion decoder estimates sparse se(3) transforms and motion base weights for dense SE(3) projection.

Quantitative Results

Table 1: Video Depth Estimation

AbsRel (↓) and δ<1.25 (↑).

MethodBonnSintelKITTI
AbsRelδAbsRelδAbsRelδ
MonST3R0.07295.480.30954.230.09591.59
DepthAnythingV30.04697.380.18670.660.05297.35
SM4RT (Ours)0.05497.490.16975.510.05697.20

Table 2: 3D Reconstruction (Pose/Rec.)

ModelHiRoomETH3DDTU
PoseRec.PoseRec.PoseRec.
Pi366.88/94.7867.6835.30/87.3071.6162.54/94.823.388
4RC85.53/98.1386.5337.79/89.9771.9792.07/99.201.380
SM4RT (Ours)86.16/98.3689.8943.88/91.5674.8693.30/99.321.153

Table 3: World-Coordinate Tracking

Average Percent of Points within Delta (APD) and End-Point Error (EPE), reported as All / Dynamic. Predictions are aligned using Global Alignment over 64 frames.

Model ADT DS PO PStudio
APDEPE APDEPE APDEPE APDEPE
MonST3R 74.35 / 67.920.2721 / 0.1578 58.06 / 51.860.4387 / 0.5313 33.47 / 39.360.9021 / 0.6452 51.32 / 51.320.4568 / 0.4568
SpatialTracker 45.65 / 67.650.8530 / 0.1628 54.85 / 58.650.9274 / 1.0828 38.54 / 51.200.7499 / 0.4695 62.59 / 62.590.3094 / 0.3094
St4RTrack 76.00 / 75.340.2680 / 0.1212 73.74 / 68.130.2682 / 0.2961 67.94 / 68.710.3140 / 0.2970 69.67 / 69.670.2637 / 0.2637
TraceAnything 76.08 / 73.050.2469 / 0.1268 60.66 / 61.230.5756 / 0.4422 39.84 / 47.391.0595 / 0.7224 71.32 / 71.320.2727 / 0.2727
Any4D 58.65 / 70.660.4048 / 0.1313 79.08 / 68.810.2203 / 0.3112 66.20 / 66.550.3741 / 0.3991 64.44 / 64.440.2977 / 0.2977
VDPM 85.96 / 72.100.1666 / 0.1359 74.27 / 77.660.2548 / 0.2123 81.10 / 82.390.1907 / 0.1708 78.73 / 78.730.1812 / 0.1812
4RC 83.22 / 72.270.1905 / 0.1422 82.17 / 73.920.1946 / 0.2747 79.35 / 75.940.2762 / 0.3517 70.91 / 70.910.2443 / 0.2433
(Open)D4RT 72.20 / 73.250.2758 / 0.3199 72.48 / 74.880.2959 / 0.2755 67.99 / 74.250.3178 / 0.2593 79.60 / 79.600.1753 / 0.1753
SM4RT (Ours) 88.56 / 80.680.1339 / 0.0935 82.61 / 71.620.1918 / 0.2912 81.62 / 73.680.2578 / 0.3808 78.69 / 78.690.1790 / 0.1790

More Visualizations

Motion Base Visualization

Check whether regions with similar colors share similar kinematics.

gif1
gif2
gif3
gif4
gif5
gif6
png1
png2
png3
png4
png5
png6

More Comparisons

SM4RT Architecture

Figure 4: Static Comparison of Tracklines. Our method obtains more structured motion compared with others, especially in background staticness.

Citation

@article{sm4rt,
  title={SM4RT: Learning Structured Motion Geometry for 4D Reconstruction},
  author={Lin, Shing Ho J. and Zheng, Wenzhao and Zhuo, Dong and Wu, Yuqi and Zhou, Jie and Lu, Jiwen},
  journal={arXiv preprint: arXiv 2607.},
  year={2026}
}