Dita — libero_spatial_no_noops

This repository is a strict LeRobot conversion of libero_spatial.pth from the original Dita implementation. It is configured for the libero_spatial_no_noops action distribution and is not interchangeable with another LIBERO suite.

Provenance and compatibility contract

  • Source checkpoint: model_ckpt/libero_spatial.pth
  • Source checkpoint SHA256: a4599b05cab634db60d593f1286f0da1f3ea84fbe7cbde83f8586761d68eb6d7
  • Conversion code commit: f506ceccb9b2816ffebb024b14aeaa2b34fbd46b
  • Dataset statistics: /home/shared/lyanam/huggingface/hub/datasets--openvla--modified_libero_rlds/snapshots/6ce6aaaaabdbe590b1eef5cd29c0d33f14a08551/libero_spatial_no_noops/1.0.0/dataset_statistics_61d1bed72aa63659b5737b75e26b670e3fb9124ed6f2e902565eed7ec09b3fb3.json
  • Dataset statistics SHA256: 26c9f341ad6304b2c18aa62154d98be6e5c193cffacfedd0b9e9b62e619a8776
  • Image processing: one RGB primary/agentview camera, resized to 224×224, followed once by ImageNet mean (0.485, 0.456, 0.406) and std (0.229, 0.224, 0.225).
  • Action processing: dimensions 0–5 use this suite's q01/q99; dimension 6 is not normalized.
  • Action mask: [true, true, true, true, true, true, false].
  • q01: [-0.7454732114076613, -0.6616071462631226, -0.9375, -0.1071428582072258, -0.20678570866584778, -0.1842857152223587, 0.0]
  • q99: [0.9375, 0.8758928775787354, 0.9321428537368774, 0.1039285734295845, 0.17678570747375488, 0.14571428298950195, 1.0]
  • Gripper: normalized output is thresholded with > 0.5, then the Cartesian dimensions are unnormalized, then the original Dita gripper sign flip is applied.
  • Rollout: traj_length=11, two-frame history, chunk_size=10, n_action_steps=1, num_inference_steps=10, DDIM prediction_type=epsilon, and CUDA AMP by default.
  • policy.dtype controls compute precision (float32, float16, bfloat16); parameters remain FP32. dtype is the sole precision setting; use_amp is derived internally.
  • LIBERO's LeRobot environment processor performs the 180° image rotation. The policy does not rotate the image a second time.

Install

pip install "lerobot_policy_dita[libero] @ https://huggingface.co/albus2024/dita_spatial/resolve/main/lerobot_policy_dita-2.2.0-py3-none-any.whl"

Evaluate

lerobot-eval \
  --policy.path=albus2024/dita_spatial \
  --policy.gripper_mode=binary \
  --policy.device=cuda \
  --env.type=libero \
  --env.task=libero_spatial \
  --env.max_parallel_tasks=1 \
  --eval.batch_size=1 \
  --eval.n_episodes=10 \
  --output_dir=./outputs/dita_spatial

The processor state in this repository contains the required statistics; no source .pth file or local dataset metadata is needed at evaluation time.

Validation

The converter checks all parameter names and shapes, loads with strict=True, and checks the saved artifact files. Run the project's parity and converted-checkpoint verification scripts separately before using the policy; conversion alone does not establish rollout parity.

Runtime 2.1.0 training configuration

Install the 2.2.0 wheel above for this revision. The new training fields require this runtime. The policy weights, processor files and action statistics are preserved; this is a configuration/runtime update, not a newly trained model.

Setting Value
freeze_backbone false
gradient_checkpointing true, only during gradient-enabled training
optimizer_lr 0.0001
optimizer_backbone_lr_scale 0.1
AdamW betas / epsilon [0.9, 0.95] / 1e-8
Weight decay 0.05; bias and one-dimensional parameters use 0
Gradient clipping Disabled (optimizer_grad_clip_norm=0)
LR schedule 1000-step warmup, cosine to 100000 steps, minimum scale 0.01
lora_dropout 0.0

Inference automatically skips ViT and QFormer activation-checkpoint wrappers. policy.dtype remains float16 by default; BF16 and FP32 are also supported. CLIP text encoding remains frozen. The gripper mode remains binary; see gripper modes.

Fine-tuning and LoRA

Use the following for full fine-tuning (choose an appropriate target dataset):

lerobot-train \
  --policy.path=albus2024/dita_spatial \
  --policy.device=cuda \
  --policy.dtype=bfloat16 \
  --policy.input_features=null \
  --policy.push_to_hub=false \
  --dataset.repo_id=YOUR_USERNAME/YOUR_DATASET \
  --output_dir=outputs/train/dita_spatial_full \
  --steps=100000

For LoRA, install the optional dependency and add the adapter arguments:

pip install "lerobot_policy_dita[lora] @ https://huggingface.co/albus2024/dita_spatial/resolve/main/lerobot_policy_dita-2.2.0-py3-none-any.whl"
# Add to lerobot-train and choose a fresh output directory:
# --peft.method_type=LORA --peft.r=32

LoRA targets every Dita linear layer, including the ViT, QFormer/FiLM and action head. Only adapter weights are trained. The default alpha is twice the rank; initialization is Gaussian for A and zero for B. Use LeRobot's native policy factory/evaluation entry points to load saved adapters together with their base model. Do not set policy.use_peft=true when starting from a base checkpoint.

The LR schedule keeps its configured 100000-step period even in shorter runs; the first update has zero LR. Set policy.scheduler_decay_steps explicitly for a different period. These defaults follow the upstream LIBERO recipe, including the higher 5e-4 peak LR for LIBERO-Long; DROID uses 1e-4 as a fine-tuning starting point and can be overridden for the target task. When adapting DROID to LIBERO, pass --policy.gripper_mode=binary.

See runtime validation for the checks performed with this release and source hashes for the exact packaged code. The previous runtime 2.0.1 report is archived. Short training/inference checks do not establish convergence or robot task success.

Runtime 2.2.0 image preprocessing

Install the 2.2.0 wheel linked above before loading this revision. The saved config now explicitly contains image_preprocessing_device=cpu. The weights, processor files and action statistics remain the same.

Pixel validation always uses one min/max reduction on the model device: on CUDA it runs on the GPU even when image resizing runs on the CPU. NaN, infinities and pixels outside [0, 1] are rejected. There is no validation-device option.

policy.image_preprocessing_device controls FP32 antialiased bilinear resizing and ImageNet normalization in both training and inference:

Value Behavior
cpu CPU resize/normalization, preserving the reference numerical path
policy Resize/normalization on the actual model device, including local CUDA ranks

To fine-tune using GPU preprocessing, add this CLI override:

--policy.image_preprocessing_device=policy

This overrides an initialization saved with cpu. The fine-tuned full-policy or LoRA checkpoint saves policy, which is restored for inference. Use the same saved setting for training and deployment. policy.dtype controls model AMP; image preprocessing remains FP32 for all supported compute dtypes.

With fixed RNG and deterministic kernels, the new cpu mode exactly matched the previous implementation's loss, all parameter gradients, and inference actions on the checked real-data batch in FP32, BF16 and FP16. Switching existing weights to GPU resizing produces measurable differences, especially under AMP; no rollout-equivalence or convergence claim is made. This published configuration therefore remains cpu.

See per-model numerical validation, training and deployment instructions, and the archived runtime 2.1.0 configuration.

Downloads last month
78
Safetensors
Model size
0.2B params
Tensor type
F32
·
Video Preview
loading

Collection including albus2024/dita_spatial