LingBot-VA Dual-Franka Five-Task LoRA — 10k

English | 中文说明

This repository contains a LoRA adapter trained from robbyant/lingbot-va-base on five real-robot manipulation tasks collected with a dual Franka FR3 setup.

The repository is not a standalone full model. Load the base LingBot-VA checkpoint first, then apply adapter_model.safetensors using adapter_config.json.

Tasks and prompts

Task ID Training prompt Episodes used
place_bowl_on_cabinet Place the green bowl on top of the cabinet with the left arm. 49
place_bowl_on_plate_pt2 Place the green bowl on the blue plate with the left arm. 18
place_bowl_on_plate Place the gray bowl on the green plate with the left arm. 49
place_cup_on_tray Place the cup on the tray with the left arm. 50
tape_on_basket Place the tape in the green basket with the left arm. 50

Total training episodes: 216. The dataset is imbalanced because only 18 place_bowl_on_plate_pt2 trajectories passed the strict LingBot-VA conversion/alignment checks.

Inputs and outputs

  • Visual inputs: chest, left-wrist, and right-wrist RGB streams
  • Language conditioning: exact task prompt listed above
  • Training representation: video features pre-encoded as Wan2.2 VAE visual latents, with text embeddings
  • Robot-state input: no observation.state was supplied to this LingBot-VA training run
  • Source action: 16 dimensions — left 7 joints + left gripper + right 7 joints + right gripper
  • LingBot action layout: the 16 source dimensions are scattered into a 30-dimensional canonical layout; unused Cartesian channels are masked/padded

Training configuration

Setting Value
Base model robbyant/lingbot-va-base
Method LoRA SFT
Trainable parameters 47,185,920
LoRA rank / alpha / dropout 32 / 32 / 0.05
Target modules self- and cross-attention Q/K/V/output projections
Training steps 10,000
Hardware 4 × NVIDIA A100 80GB
Per-device / global batch 1 / 4
Gradient accumulation 1
Learning rate 5e-5, 10-step warmup
Adam betas / weight decay 0.9, 0.95 / 0.1
Precision bfloat16
Checkpoint interval 1,000 steps
Final checkpoint checkpoint_step_10000

At step 9,999, the logged losses were: latent 0.0806, action 0.00481, total 0.0854. These are training losses, not task success rates.

Repository contents

  • adapter_model.safetensors: LoRA weights
  • adapter_config.json: rank, alpha, dropout, target modules, and base-model reference
  • training_state.pt: optimizer/scheduler/RNG state for the original training environment

training_state.pt may contain Python-serialized objects. Only load it with torch.load if you trust this repository and need training resumption. It is not required for inference.

Download and loading

hf download robbyant/lingbot-va-base --local-dir ./lingbot-va-base
hf download THU98/lingbotva_franka_5task_lora_10k \
  --local-dir ./lingbotva_franka_5task_lora_10k

Use the official LingBot-VA repository and its post-training/inference code:

git clone https://github.com/Robbyant/lingbot-va.git

Configure the base path as ./lingbot-va-base and the adapter path as:

./lingbotva_franka_5task_lora_10k/adapter_model.safetensors

The exact adapter targets are stored in adapter_config.json. Use the dual-Franka joint configuration and the same 16-to-30-dimensional action-channel mapping used during training. Generic PEFT loading is not guaranteed because the training code uses the LingBot-VA LoRA integration.

Evaluation status

No controlled real-robot success-rate evaluation is published with this adapter. No held-out validation set was used to select the final checkpoint. Loss reduction alone does not demonstrate safe or successful robot execution.

Intended use and limitations

  • Research use with the official LingBot-VA base model and a matching dual Franka setup.
  • The data are small, task-specific, imbalanced, and dominated by left-arm behaviors.
  • The adapter was trained without proprioceptive observation.state input; behavior depends heavily on visual calibration and temporal alignment.
  • It is not validated for different robots, viewpoints, objects, control frequencies, action conventions, or prompt paraphrases.
  • Validate action-channel mapping, normalization, camera order, control rate, latency, collision avoidance, and emergency stops before deployment.
  • Human supervision and independent robot safety controls are required.

中文说明

这是基于 robbyant/lingbot-va-base 训练的双 Franka FR3 五任务 LoRA adapter,不是可以独立加载的完整模型。推理时需要先加载 LingBot-VA base,再加载本仓库中的 adapter_model.safetensors

训练使用三路相机的视觉 latent 和 Language Instruction,没有把 16 维 observation.state 作为模型输入。源动作是 16 维双臂关节与夹爪动作,并映射到 LingBot-VA 的 30 维动作布局。五个任务共使用 216 条轨迹,其中 place_bowl_on_plate_pt2 只有 18 条通过严格数据校验,因此数据分布不均衡。

当前没有发布正式真机成功率。部署前必须核对 base 版本、LoRA target、相机顺序、动作通道、归一化、控制频率以及真机安全限制。

Sources and citation

@article{lingbotva2026,
  title={Causal World Modeling for Robot Control},
  journal={arXiv preprint arXiv:2601.21998},
  year={2026}
}
Downloads last month
6
Video Preview
loading

Model tree for THU98/lingbotva_franka_5task_lora_10k

Adapter
(1)
this model

Paper for THU98/lingbotva_franka_5task_lora_10k