Mistral’s Robostral Navigate lets robots move using a single camera
Mistral released Robostral Navigate, an 8B embodied model that guides robots through complex spaces from one RGB camera and plain-language prompts.
Mistral AI has stepped further into physical AI with Robostral Navigate, an 8-billion-parameter model that lets robots find their way through unfamiliar spaces using nothing more than a single camera and natural-language instructions. Announced on 8 July 2026, the company calls it its first model built for embodied navigation.
The design deliberately avoids expensive sensor stacks. Instead of LiDAR, depth cameras or multi-camera rigs, Robostral Navigate works from one ordinary RGB camera. It uses what Mistral describes as a pointing-based method: the model predicts the image coordinates of a target location in the robot’s current view, along with a desired orientation. When a destination sits outside the camera’s field of view, it falls back to local displacement instructions such as moving a set distance forward and to the side. The approach is built on a vision-language model specialised in grounding tasks, with navigation emerging as an extension of that ability.
Because it reads generic camera input, the model is hardware-agnostic and works across wheeled, legged and flying robots of different sizes. Mistral says it can handle offices, homes, commercial buildings and outdoor settings while avoiding obstacles.
Training took place entirely in simulation, using roughly 400,000 trajectories across 6,000 scenes. Mistral reports an efficiency gain from prefix caching that cut training tokens by a factor of 22, followed by an online reinforcement-learning stage using the CISPO algorithm that lifted success rates by a further 3.2 percent. On evaluation, the model reached a 79.4 percent success rate in seen environments and 76.6 percent on the unseen R2R-CE benchmark, which the company says beats the best single-camera approach by 9.7 points and multi-sensor systems by 4.5 points.
The release signals how quickly established language-model labs are pivoting toward robotics, where the promise of low-cost, camera-only navigation could widen the range of machines that can operate autonomously without costly instrumentation.
Sources: Mistral AI
Drafted with AI assistance and reviewed by the IT Magazine news desk. Facts verified against the linked sources. Spotted an error? See our Corrections Policy.



