Alpamayo1.5-10B β Renesas X5H
βΆ NVIDIA Alpamayo β Thinking Out Loud β official demo video, via nvidia.com
π§ Coming soon. This model is being optimized and validated for Renesas hardware β no fixed release date yet.
Introduction
Alpamayo1.5-10B is NVIDIA's vision-language-action (VLA) model for autonomous-vehicle research, built for researchers and driving practitioners working on rare, long-tail driving scenarios. It combines a Cosmos-Reason2 VLM backbone (8.2B params) with a diffusion-based trajectory decoder / action expert (2.3B params) β ~10.5B params total. Renesas is preparing an optimized deployment of this model for the R-Car Gen5 platform.
Unlike general-purpose robotics VLAs (e.g. GR00T), Alpamayo is scoped specifically to driving: it takes multi-camera RGB video (4 cameras @ 10Hz), text commands/navigation guidance, and egomotion history as input, and outputs chain-of-thought reasoning explaining its driving decisions plus a 6.4-second trajectory prediction.
- Model Architecture: Vision-language-action (VLA) β Cosmos-Reason2 VLM backbone + diffusion trajectory decoder, ~10.5B parameters.
- Source Model: nvidia/Alpamayo-1.5-10B
This page is a placeholder. Content is provisional and subject to change before the model is fully published. Upstream weights are under NVIDIA's OpenMDW-1.1 license (non-commercial; commercial licensing available on request); upstream source code is Apache-2.0.
Model tree for Renesas/Alpamayo1.5-10B
Base model
Qwen/Qwen3-VL-8B-Instruct