TMCnet News

Alibaba's Amap Introduces ABot-Recon, Reconstructing 10,000-Frame-Scale 3D Scenes From Just 12 Frames in Real Time
[August 28, 2026]

Alibaba's Amap Introduces ABot-Recon, Reconstructing 10,000-Frame-Scale 3D Scenes From Just 12 Frames in Real Time


BEIJING, Aug. 28, 2026 /PRNewswire/ -- Amap, Alibaba's location-based services platform, today released ABot-Recon, a streaming 3D reconstruction model that requires only 12 consecutive frames to reconstruct scenes spanning over 10,000 frames in real time, eliminating the need for long-range memory.

The model achieves state-of-the-art accuracy on multiple public benchmarks — including KITTI, Oxford Spires and VBR — while reducing peak memory usage to roughly one-third that of comparable methods, making real-time 3D reconstruction feasible on consumer-grade hardware.

As autonomous driving and embodied AI systems move into open environments, they require continuous spatial understanding while in motion — knowing where they are, what surrounds them and where to go next. Conventional streaming 3D reconstruction systems maintain memory anchors that store and fuse historical information to preserve global consistency. As input sequences grow longer, such systems become progressively slower, less accurate and more memory-intensive.

ABot-Recon takes a fundamentally different approach. Rther than maintaining long-range memory anchors, the model operates within a fixed 12-frame local context window, predicting only the local point cloud and the relative pose between adjacent frames. An online composition mechanism then assembles the complete global trajectory incrementally, keeping computational complexity constant regardless of sequence length.

To address drift inherent in local prediction, ABot-Recon incorporates dedicated correction and constraint mechanisms at both the prediction and training stages, calibrating trajectory error in real time.

On the Oxford Spires long-sequence benchmark, ABot-Recon reduces average trajectory error by 40.6% compared with the prior leading method, achieving a relative rotation error (RPE-R) of 0.12 degrees — approximately 40% lower than the previous state of the art. On KITTI-02, the model achieves real-time reconstruction at 24.45 FPS, 1.24 times the speed of existing approaches, with peak memory usage of approximately 6.71 GB — meaning a consumer-grade GTX 1080 Ti is sufficient to run the full pipeline.


The model requires only monocular RGB video as input, and needs no depth sensors or pre-calibrated camera parameters. This positions ABot-Recon for deployment across private-area mapping, embodied AI training, autonomous driving and 3D content production — scenarios where pre-built maps are unavailable and real-time reconstruction is essential.

ABot-Recon's inference code, evaluation scripts and pre-trained weights are now open-sourced on GitHub. The project page is accessible at https://amap-cvlab.github.io/ABot-Recon-html/.

Cision View original content to download multimedia:https://www.prnewswire.com/news-releases/alibabas-amap-introduces-abot-recon-reconstructing-10-000-frame-scale-3d-scenes-from-just-12-frames-in-real-time-302862527.html

SOURCE Amap


[ Back To TMCnet.com's Homepage ]