DM05-SO101-Pick-Cube

DM0.5

Tech Blog GitHub MaaS

Introduction

DM05-SO101-Pick-Cube is the SO101 fine-tuned checkpoint of DM0.5, Dexmal's open-world Vision-Language-Action foundation model for embodied intelligence. DM0.5 uses a Gemma3 4B vision-language backbone with a 680M Action Expert to generate continuous robot actions, and is designed for natural-language manipulation, zero-shot generalization, efficient downstream fine-tuning, long-horizon historical context, robust policy behavior, and transfer across robot embodiments.

This checkpoint is specifically trained for the SO101 pick cube task using LoRA fine-tuning.

Quick Start

We recommend using Docker to set up the runtime environment first, which helps avoid version mismatches across CUDA, PyTorch, flash-attn, and other dependencies on the host machine.

Requirements

System requirements:
Ubuntu 20.04 / 22.04
NVIDIA GPU
NVIDIA Driver
Docker
NVIDIA Container Toolkit
Conda (optional, only required for local pip installation)

Recommended GPUs:
RTX 4090, A100, H100, H20
8 GPUs are recommended for training, and 1 GPU is sufficient for deployment inference.

Docker Installation

git clone https://github.com/dexmal/opendm.git
cd opendm

docker run -it --rm --gpus all --network host \
  --name opendm \
  --shm-size=16g \
  -v "$PWD":/app/opendm \
  -w /app/opendm \
  dexmal/opendm:latest /bin/bash

# Run from the OpenDM repository root inside the container.
conda activate opendm
pip install -e .

Local Installation

conda create -n opendm python=3.10 -y
conda activate opendm

pip install torch torchvision \
  --index-url https://download.pytorch.org/whl/cu128

pip install ninja packaging
MAX_JOBS=2 pip install flash-attn --no-build-isolation

# Enter the OpenDM repository root.
cd opendm
pip install -e .

SO101 Inference

Use the SO101-specific experiment configuration when running inference with this checkpoint. Run this command from the OpenDM repository root:

script/dm05_launcher.sh \
  --exp playground/dm05_so101_lora.py \
  --task inference \
  --nproc_per_node 1 \
  --model-config.model-name-or-path ./checkpoints/DM05-SO101-Pick-Cube \
  --model-config.chunk-size 50 \
  --inference-config.output-action-dim 6 \
  --inference-config.image-keys images_1 images_2 \
  --inference-config.port 7891

For the complete training and inference workflow, see the DM05 SO101 LoRA Training Guide.

Community and Support

  • Learn more about Dexmal products and model updates on the Dexmal website.
  • If you encounter issues, please report them through GitHub Issues.
  • For further discussion, scan the WeChat QR code to contact us.

We will continue to release more model weights, technical documentation, and examples. If this project is helpful to you, please consider giving us a star on GitHub GitHub. Your support helps us move forward.

Citation

@misc{dm05,
    title  = {{DM0.5}: An Open-World Foundation Model for General-Purpose Embodied Intelligence},
    author = {{Dexmal Team}},
    month  = {July},
    year   = {2026},
    url    = {https://www.dexmal.com/blog/dm0.5/index_en.html}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for Dexmal/DM05-SO101-Pick-Cube

Base model

Dexmal/DM05
Finetuned
(3)
this model

Collection including Dexmal/DM05-SO101-Pick-Cube