Alpamayo1.5-10B – Renesas X5H

Watch: NVIDIA Alpamayo – Thinking Out Loud

β–Ά NVIDIA Alpamayo – Thinking Out Loud β€” official demo video, via nvidia.com

🚧 Coming soon. This model is being optimized and validated for Renesas hardware β€” no fixed release date yet.

Introduction

Alpamayo1.5-10B is NVIDIA's vision-language-action (VLA) model for autonomous-vehicle research, built for researchers and driving practitioners working on rare, long-tail driving scenarios. It combines a Cosmos-Reason2 VLM backbone (8.2B params) with a diffusion-based trajectory decoder / action expert (2.3B params) β€” ~10.5B params total. Renesas is preparing an optimized deployment of this model for the R-Car Gen5 platform.

Unlike general-purpose robotics VLAs (e.g. GR00T), Alpamayo is scoped specifically to driving: it takes multi-camera RGB video (4 cameras @ 10Hz), text commands/navigation guidance, and egomotion history as input, and outputs chain-of-thought reasoning explaining its driving decisions plus a 6.4-second trajectory prediction.

  • Model Architecture: Vision-language-action (VLA) β€” Cosmos-Reason2 VLM backbone + diffusion trajectory decoder, ~10.5B parameters.
  • Source Model: nvidia/Alpamayo-1.5-10B

This page is a placeholder. Content is provisional and subject to change before the model is fully published. Upstream weights are under NVIDIA's OpenMDW-1.1 license (non-commercial; commercial licensing available on request); upstream source code is Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for Renesas/Alpamayo1.5-10B

Finetuned
(2)
this model