Instructions to use albus2024/dita_droid_pretrain with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use albus2024/dita_droid_pretrain with LeRobot:
- Notebooks
- Google Colab
- Kaggle
Dita DROID pretraining initialization
This repository converts the original Dita DROID weights for subsequent fine-tuning. It contains model weights, a LeRobot Dita configuration, and processor templates without dataset action statistics. It is not a ready-to-deploy DROID or LIBERO policy.
Provenance
- Source:
model_ckpt/dita_droid_pretrain.pthfrom the original Dita release. - Source checkpoint SHA256:
30c5056d61d68be67ea509458b152dd9130b1c9bdc046e27b2157ce2f63220af. - Conversion base Git commit:
c97999bb495b3d58b105772690138a2cb7e8a709. Seeconversion_metadata.jsonfor the working-tree status, conversion-script hash, and packaged runtime hash when present. - All model parameter keys and shapes are checked, followed by
strict=Trueloading. - Optimizer/scheduler state is not exported. No source DROID action statistics are inferred.
Install
Use Python 3.12. The updated Dita runtime translates LeRobot's standard fine-tuning normalizer overrides to Dita's custom processors.
pip install "lerobot_policy_dita[training] @ https://huggingface.co/albus2024/dita_droid_pretrain/resolve/main/lerobot_policy_dita-2.2.0-py3-none-any.whl"
Fine-tune
Use a LeRobot dataset with seven-dimensional actions, a compatible primary RGB camera,
and task language. Its action metadata must contain seven-element q01 and q99 arrays
computed in the target training action convention. The default gripper_mode=continuous
uses quantiles for the first six dimensions and dataset action.min/max for the gripper.
Continuous outputs are restored to the recorded position units without thresholding or
sign inversion. Use --policy.gripper_mode=binary for LIBERO signed training targets
(-1=open, +1=closed), encoded internally as 1=open, 0=closed.
ImageNet image normalization and CLIP text encoding remain fixed.
lerobot-train \
--policy.path=albus2024/dita_droid_pretrain \
--policy.gripper_mode=continuous \
--policy.device=cuda \
--policy.dtype=bfloat16 \
--policy.input_features=null \
--policy.push_to_hub=false \
--dataset.repo_id=YOUR_USERNAME/YOUR_DATASET \
--output_dir=outputs/train/dita_droid_finetune
policy.dtype accepts float32, float16 and bfloat16 and controls autocast;
weights and optimizer state remain FP32. Use BF16 on a supported GPU, or select
FP32/FP16 explicitly. The default CUDA precision remains FP16 for compatibility.
At the start of a new fine-tuning run, LeRobot binds the target dataset's statistics to both processor templates. They remain fixed during that run and are saved with the fine-tuned checkpoint. LeRobot 0.6.1 restores saved statistics on resume. The same pretraining repository can initialize multiple target datasets; no per-dataset pretraining upload is needed. Missing quantiles cause an explicit error when processing actions.
The processor templates require cached or downloadable DINOv2/CLIP assets. Load the
fine-tuned checkpoint together with its saved processors for evaluation and deployment.
Any validation_report.json documents the checks actually run for this upload; no
task success rate or DROID deployment equivalence is claimed.
Runtime 2.1.0 training configuration
Install the 2.2.0 wheel above for this revision. The new training fields require this runtime. The policy weights, processor files and action statistics are preserved; this is a configuration/runtime update, not a newly trained model.
| Setting | Value |
|---|---|
freeze_backbone |
false |
gradient_checkpointing |
true, only during gradient-enabled training |
optimizer_lr |
0.0001 |
optimizer_backbone_lr_scale |
0.1 |
| AdamW betas / epsilon | [0.9, 0.95] / 1e-8 |
| Weight decay | 0.05; bias and one-dimensional parameters use 0 |
| Gradient clipping | Disabled (optimizer_grad_clip_norm=0) |
| LR schedule | 1000-step warmup, cosine to 100000 steps, minimum scale 0.01 |
lora_dropout |
0.0 |
Inference automatically skips ViT and QFormer activation-checkpoint wrappers.
policy.dtype remains float16 by default; BF16 and FP32 are also supported.
CLIP text encoding remains frozen. The gripper mode remains
continuous; see gripper modes.
Fine-tuning and LoRA
Use the following for full fine-tuning (choose an appropriate target dataset):
lerobot-train \
--policy.path=albus2024/dita_droid_pretrain \
--policy.device=cuda \
--policy.dtype=bfloat16 \
--policy.input_features=null \
--policy.push_to_hub=false \
--dataset.repo_id=YOUR_USERNAME/YOUR_DATASET \
--output_dir=outputs/train/dita_droid_pretrain_full \
--steps=100000
For LoRA, install the optional dependency and add the adapter arguments:
pip install "lerobot_policy_dita[lora] @ https://huggingface.co/albus2024/dita_droid_pretrain/resolve/main/lerobot_policy_dita-2.2.0-py3-none-any.whl"
# Add to lerobot-train and choose a fresh output directory:
# --peft.method_type=LORA --peft.r=32
LoRA targets every Dita linear layer, including the ViT, QFormer/FiLM and action
head. Only adapter weights are trained. The default alpha is twice the rank;
initialization is Gaussian for A and zero for B. Use LeRobot's native policy
factory/evaluation entry points to load saved adapters together with their base
model. Do not set policy.use_peft=true when starting from a base checkpoint.
The LR schedule keeps its configured 100000-step period even in shorter runs;
the first update has zero LR. Set policy.scheduler_decay_steps explicitly for
a different period. These defaults follow the upstream LIBERO recipe, including
the higher 5e-4 peak LR for LIBERO-Long; DROID uses 1e-4 as a fine-tuning
starting point and can be overridden for the target task. When adapting DROID
to LIBERO, pass --policy.gripper_mode=binary.
See runtime validation for the checks performed with this release and source hashes for the exact packaged code. The previous runtime 2.0.1 report is archived. Short training/inference checks do not establish convergence or robot task success.
Runtime 2.2.0 image preprocessing
Install the 2.2.0 wheel linked above before loading this revision. The saved
config now explicitly contains image_preprocessing_device=cpu. The weights,
processor files and action statistics remain the same.
Pixel validation always uses one min/max reduction on the model device: on CUDA
it runs on the GPU even when image resizing runs on the CPU. NaN, infinities and
pixels outside [0, 1] are rejected. There is no validation-device option.
policy.image_preprocessing_device controls FP32 antialiased bilinear resizing
and ImageNet normalization in both training and inference:
| Value | Behavior |
|---|---|
cpu |
CPU resize/normalization, preserving the reference numerical path |
policy |
Resize/normalization on the actual model device, including local CUDA ranks |
To fine-tune using GPU preprocessing, add this CLI override:
--policy.image_preprocessing_device=policy
This overrides an initialization saved with cpu. The fine-tuned full-policy
or LoRA checkpoint saves policy, which is restored for inference. Use the same
saved setting for training and deployment. policy.dtype controls model AMP;
image preprocessing remains FP32 for all supported compute dtypes.
With fixed RNG and deterministic kernels, the new cpu mode exactly matched the
previous implementation's loss, all parameter gradients, and inference actions
on the checked real-data batch in FP32, BF16 and FP16. Switching existing weights
to GPU resizing produces measurable differences, especially under AMP; no
rollout-equivalence or convergence claim is made. This published configuration
therefore remains cpu.
See per-model numerical validation, training and deployment instructions, and the archived runtime 2.1.0 configuration.
- Downloads last month
- 129