lingbot-map

Research code for a transformer model that reconstructs 3D scenes from streaming video. It accompanies an academic paper and targets 3D mapping research.

Share on XLicense: Apache-2.0

Overview

LingBot-Map is a feed-forward 3D foundation model for streaming 3D reconstruction, from the Robbyant Team. It uses a Geometric Context Transformer that combines coordinate grounding, dense geometric cues and long-range drift correction through anchor context, a pose-reference window and trajectory memory. The README reports around 20 FPS at 518 by 378 resolution over sequences beyond 10,000 frames. The repository has an interactive demo, windowed inference for long sequences, sky masking and an offline rendering pipeline.

Key features

  • Streaming 3D reconstruction with a Geometric Context Transformer
  • Paged KV cache attention, about 20 FPS at 518x378
  • Windowed inference for sequences over 3000 frames
  • Interactive demo with sky masking and visualization options
  • Evaluation scripts for KITTI and Oxford Spires

Best for

Researchers and engineers working on long-video 3D reconstruction who want a released model with demo and evaluation code. The project is presented as an ECCV 2026 oral.

Upstream
Robbyant/lingbot-map
Fork on GitHub
Guo-astro/lingbot-map
Upstream stars
17k
Category
Finance, data and research
Language
Python
License
Apache-2.0
Forked
2026-07-19
Sync status
In syncLast synced 2026-09-29