Robbyant/lingbot-map

A feed-forward 3D foundation model for reconstructing scenes from streaming data
LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction
Robbyant Teamgithub.com/user-attachments/assets/fe39e095-af2c-4ec9-b68d-a8ba97e505ab
🗺️ Meet LingBot-Map! We've built a feed-forward 3D foundation model for streaming 3D reconstruction! 🏗️🌍
LingBot-Map has focused on:
Geometric Context Transformer: Architecturally unifies coordinate grounding, dense geometric cues, and long-range drift correction within a single streaming framework through anchor context, pose-reference window, and trajectory memory.
High-Efficiency Streaming Inference: A feed-forward architecture with paged KV cache attention, enabling stable inference at ~20 FPS on 518×378 resolution over long sequenc…