Dita DROID pretraining initialization

This repository converts the original Dita DROID weights for subsequent fine-tuning. It contains model weights, a LeRobot Dita configuration, and processor templates without dataset action statistics. It is not a ready-to-deploy DROID or LIBERO policy.

Provenance

  • Source: model_ckpt/dita_droid_pretrain.pth from the original Dita release.
  • Source checkpoint SHA256: 30c5056d61d68be67ea509458b152dd9130b1c9bdc046e27b2157ce2f63220af.
  • Conversion base Git commit: c97999bb495b3d58b105772690138a2cb7e8a709. See conversion_metadata.json for the working-tree status, conversion-script hash, and packaged runtime hash when present.
  • All model parameter keys and shapes are checked, followed by strict=True loading.
  • Optimizer/scheduler state is not exported. No source DROID action statistics are inferred.

Install

Use Python 3.12. The updated Dita runtime translates LeRobot's standard fine-tuning normalizer overrides to Dita's custom processors.

pip install "lerobot_policy_dita[training] @ https://huggingface.co/albus2024/dita_droid_pretrain/resolve/main/lerobot_policy_dita-2.2.0-py3-none-any.whl"

Fine-tune

Use a LeRobot dataset with seven-dimensional actions, a compatible primary RGB camera, and task language. Its action metadata must contain seven-element q01 and q99 arrays computed in the target training action convention. The default gripper_mode=continuous uses quantiles for the first six dimensions and dataset action.min/max for the gripper. Continuous outputs are restored to the recorded position units without thresholding or sign inversion. Use --policy.gripper_mode=binary for LIBERO signed training targets (-1=open, +1=closed), encoded internally as 1=open, 0=closed. ImageNet image normalization and CLIP text encoding remain fixed.

lerobot-train \
  --policy.path=albus2024/dita_droid_pretrain \
  --policy.gripper_mode=continuous \
  --policy.device=cuda \
  --policy.dtype=bfloat16 \
  --policy.input_features=null \
  --policy.push_to_hub=false \
  --dataset.repo_id=YOUR_USERNAME/YOUR_DATASET \
  --output_dir=outputs/train/dita_droid_finetune

policy.dtype accepts float32, float16 and bfloat16 and controls autocast; weights and optimizer state remain FP32. Use BF16 on a supported GPU, or select FP32/FP16 explicitly. The default CUDA precision remains FP16 for compatibility.

At the start of a new fine-tuning run, LeRobot binds the target dataset's statistics to both processor templates. They remain fixed during that run and are saved with the fine-tuned checkpoint. LeRobot 0.6.1 restores saved statistics on resume. The same pretraining repository can initialize multiple target datasets; no per-dataset pretraining upload is needed. Missing quantiles cause an explicit error when processing actions.

The processor templates require cached or downloadable DINOv2/CLIP assets. Load the fine-tuned checkpoint together with its saved processors for evaluation and deployment. Any validation_report.json documents the checks actually run for this upload; no task success rate or DROID deployment equivalence is claimed.

Runtime 2.1.0 training configuration

Install the 2.2.0 wheel above for this revision. The new training fields require this runtime. The policy weights, processor files and action statistics are preserved; this is a configuration/runtime update, not a newly trained model.

Setting Value
freeze_backbone false
gradient_checkpointing true, only during gradient-enabled training
optimizer_lr 0.0001
optimizer_backbone_lr_scale 0.1
AdamW betas / epsilon [0.9, 0.95] / 1e-8
Weight decay 0.05; bias and one-dimensional parameters use 0
Gradient clipping Disabled (optimizer_grad_clip_norm=0)
LR schedule 1000-step warmup, cosine to 100000 steps, minimum scale 0.01
lora_dropout 0.0

Inference automatically skips ViT and QFormer activation-checkpoint wrappers. policy.dtype remains float16 by default; BF16 and FP32 are also supported. CLIP text encoding remains frozen. The gripper mode remains continuous; see gripper modes.

Fine-tuning and LoRA

Use the following for full fine-tuning (choose an appropriate target dataset):

lerobot-train \
  --policy.path=albus2024/dita_droid_pretrain \
  --policy.device=cuda \
  --policy.dtype=bfloat16 \
  --policy.input_features=null \
  --policy.push_to_hub=false \
  --dataset.repo_id=YOUR_USERNAME/YOUR_DATASET \
  --output_dir=outputs/train/dita_droid_pretrain_full \
  --steps=100000

For LoRA, install the optional dependency and add the adapter arguments:

pip install "lerobot_policy_dita[lora] @ https://huggingface.co/albus2024/dita_droid_pretrain/resolve/main/lerobot_policy_dita-2.2.0-py3-none-any.whl"
# Add to lerobot-train and choose a fresh output directory:
# --peft.method_type=LORA --peft.r=32

LoRA targets every Dita linear layer, including the ViT, QFormer/FiLM and action head. Only adapter weights are trained. The default alpha is twice the rank; initialization is Gaussian for A and zero for B. Use LeRobot's native policy factory/evaluation entry points to load saved adapters together with their base model. Do not set policy.use_peft=true when starting from a base checkpoint.

The LR schedule keeps its configured 100000-step period even in shorter runs; the first update has zero LR. Set policy.scheduler_decay_steps explicitly for a different period. These defaults follow the upstream LIBERO recipe, including the higher 5e-4 peak LR for LIBERO-Long; DROID uses 1e-4 as a fine-tuning starting point and can be overridden for the target task. When adapting DROID to LIBERO, pass --policy.gripper_mode=binary.

See runtime validation for the checks performed with this release and source hashes for the exact packaged code. The previous runtime 2.0.1 report is archived. Short training/inference checks do not establish convergence or robot task success.

Runtime 2.2.0 image preprocessing

Install the 2.2.0 wheel linked above before loading this revision. The saved config now explicitly contains image_preprocessing_device=cpu. The weights, processor files and action statistics remain the same.

Pixel validation always uses one min/max reduction on the model device: on CUDA it runs on the GPU even when image resizing runs on the CPU. NaN, infinities and pixels outside [0, 1] are rejected. There is no validation-device option.

policy.image_preprocessing_device controls FP32 antialiased bilinear resizing and ImageNet normalization in both training and inference:

Value Behavior
cpu CPU resize/normalization, preserving the reference numerical path
policy Resize/normalization on the actual model device, including local CUDA ranks

To fine-tune using GPU preprocessing, add this CLI override:

--policy.image_preprocessing_device=policy

This overrides an initialization saved with cpu. The fine-tuned full-policy or LoRA checkpoint saves policy, which is restored for inference. Use the same saved setting for training and deployment. policy.dtype controls model AMP; image preprocessing remains FP32 for all supported compute dtypes.

With fixed RNG and deterministic kernels, the new cpu mode exactly matched the previous implementation's loss, all parameter gradients, and inference actions on the checked real-data batch in FP32, BF16 and FP16. Switching existing weights to GPU resizing produces measurable differences, especially under AMP; no rollout-equivalence or convergence claim is made. This published configuration therefore remains cpu.

See per-model numerical validation, training and deployment instructions, and the archived runtime 2.1.0 configuration.

Downloads last month
129
Safetensors
Model size
0.2B params
Tensor type
F32
·
Video Preview
loading

Collection including albus2024/dita_droid_pretrain