Figure 1: SM4RT decomposes scene motion into structured latent bases.
Recent advances in Geometry Foundation Models (GFMs) have demonstrated remarkable capability in reconstructing 3D scene geometry from monocular video. With static geometry reconstruction now approaching practical fidelity, the next frontier is 4D dynamic understanding which jointly infers scene structure and the underlying motion.
We propose SM4RT, a Structured Motion 4D Reconstruction Transformer framework. Rather than predicting independent point-wise displacements, SM4RT introduces Structure-of-Motion (SoM): scene motion is composed of a compact set of N latent motion bases, each represented as a temporal sequence of 6D se(3) transforms.
Motivation1: Humans perceive motion via sparse kinematic cues—not at the pixel level, but at the level of objects and parts. Real-world objects usually obey rigid-body kinematics: points move collectively, not in isolation.
Motivation2: Motion exhibits geometric structure as fiber bundle, formalizing scene dynamics as a principal bundle whose local sections correspond to valid rigid-body motions, ensuring that point-wise displacements respect the underlying topological and kinematic constraints.
Figure 2: Comparison of motion representations. The proposed representation preserves the structure of object motion, providing a more structured description than standard per-pixel displacement and encouraging intra-object homogeneity.
Figure 3: Architecture of SM4RT.
Given a monocular input video, SM4RT extracts latent geometry tokens with a DINOv2 backbone, then decouples them into scene geometry and motion geometry tokens. The scene decoder predicts camera parameters and depth, while the motion decoder estimates sparse se(3) transforms and motion base weights for dense SE(3) projection.
AbsRel (↓) and δ<1.25 (↑).
| Method | Bonn | Sintel | KITTI | |||
|---|---|---|---|---|---|---|
| AbsRel | δ | AbsRel | δ | AbsRel | δ | |
| MonST3R | 0.072 | 95.48 | 0.309 | 54.23 | 0.095 | 91.59 |
| DepthAnythingV3 | 0.046 | 97.38 | 0.186 | 70.66 | 0.052 | 97.35 |
| SM4RT (Ours) | 0.054 | 97.49 | 0.169 | 75.51 | 0.056 | 97.20 |
| Model | HiRoom | ETH3D | DTU | |||
|---|---|---|---|---|---|---|
| Pose | Rec. | Pose | Rec. | Pose | Rec. | |
| Pi3 | 66.88/94.78 | 67.68 | 35.30/87.30 | 71.61 | 62.54/94.82 | 3.388 |
| 4RC | 85.53/98.13 | 86.53 | 37.79/89.97 | 71.97 | 92.07/99.20 | 1.380 |
| SM4RT (Ours) | 86.16/98.36 | 89.89 | 43.88/91.56 | 74.86 | 93.30/99.32 | 1.153 |
Average Percent of Points within Delta (APD) and End-Point Error (EPE), reported as All / Dynamic. Predictions are aligned using Global Alignment over 64 frames.
| Model | ADT | DS | PO | PStudio | ||||
|---|---|---|---|---|---|---|---|---|
| APD | EPE | APD | EPE | APD | EPE | APD | EPE | |
| MonST3R | 74.35 / 67.92 | 0.2721 / 0.1578 | 58.06 / 51.86 | 0.4387 / 0.5313 | 33.47 / 39.36 | 0.9021 / 0.6452 | 51.32 / 51.32 | 0.4568 / 0.4568 |
| SpatialTracker | 45.65 / 67.65 | 0.8530 / 0.1628 | 54.85 / 58.65 | 0.9274 / 1.0828 | 38.54 / 51.20 | 0.7499 / 0.4695 | 62.59 / 62.59 | 0.3094 / 0.3094 |
| St4RTrack | 76.00 / 75.34 | 0.2680 / 0.1212 | 73.74 / 68.13 | 0.2682 / 0.2961 | 67.94 / 68.71 | 0.3140 / 0.2970 | 69.67 / 69.67 | 0.2637 / 0.2637 |
| TraceAnything | 76.08 / 73.05 | 0.2469 / 0.1268 | 60.66 / 61.23 | 0.5756 / 0.4422 | 39.84 / 47.39 | 1.0595 / 0.7224 | 71.32 / 71.32 | 0.2727 / 0.2727 |
| Any4D | 58.65 / 70.66 | 0.4048 / 0.1313 | 79.08 / 68.81 | 0.2203 / 0.3112 | 66.20 / 66.55 | 0.3741 / 0.3991 | 64.44 / 64.44 | 0.2977 / 0.2977 |
| VDPM | 85.96 / 72.10 | 0.1666 / 0.1359 | 74.27 / 77.66 | 0.2548 / 0.2123 | 81.10 / 82.39 | 0.1907 / 0.1708 | 78.73 / 78.73 | 0.1812 / 0.1812 |
| 4RC | 83.22 / 72.27 | 0.1905 / 0.1422 | 82.17 / 73.92 | 0.1946 / 0.2747 | 79.35 / 75.94 | 0.2762 / 0.3517 | 70.91 / 70.91 | 0.2443 / 0.2433 |
| (Open)D4RT | 72.20 / 73.25 | 0.2758 / 0.3199 | 72.48 / 74.88 | 0.2959 / 0.2755 | 67.99 / 74.25 | 0.3178 / 0.2593 | 79.60 / 79.60 | 0.1753 / 0.1753 |
| SM4RT (Ours) | 88.56 / 80.68 | 0.1339 / 0.0935 | 82.61 / 71.62 | 0.1918 / 0.2912 | 81.62 / 73.68 | 0.2578 / 0.3808 | 78.69 / 78.69 | 0.1790 / 0.1790 |
Check whether regions with similar colors share similar kinematics.












Figure 4: Static Comparison of Tracklines. Our method obtains more structured motion compared with others, especially in background staticness.
@article{sm4rt,
title={SM4RT: Learning Structured Motion Geometry for 4D Reconstruction},
author={Lin, Shing Ho J. and Zheng, Wenzhao and Zhuo, Dong and Wu, Yuqi and Zhou, Jie and Lu, Jiwen},
journal={arXiv preprint: arXiv 2607.},
year={2026}
}