VLAct Β· Qwen3-VL-4B PI Β· LIBERO-Plus
This repository contains the 50K-step VLAct downstream fine-tuning checkpoint for
LIBERO-Plus, introduced in Beyond Data Scaling: Representation-Centric Continued
Pre-training for Vision-Language-Action Models. It starts from
StarVLA/VLAct_Qwen3_Pretrain and adapts the
shared VLAct backbone with a randomly initialized PI (QwenPI_v4) action head for the target
benchmark.
This is a StarVLA training / evaluation checkpoint, not a standard
transformers.AutoModelpackage. Load it with the matching StarVLA framework (QwenPI_v4) and the packagedconfig.yaml/dataset_statistics.json. Safe robot deployment still requires embodiment-specific action mapping, normalization, camera calibration, control-rate handling, workspace constraints, and independent safety systems.
What is VLAct?
VLAct is a representation-centric continued-pretraining recipe for vision-language-action models. It preserves the VLM prior, co-trains multiple continuous action heads on a shared latent, and shares action semantics across embodiments with a partially unified padded layout and wrap-aware joint loss. Downstream policies discard the pretraining heads, randomly initialize a target action head, and fine-tune from the VLAct backbone.
The complete method, ablations, and evaluation protocols are documented in the paper and code repository.
Checkpoint details
| Item | Value |
|---|---|
| Framework | StarVLA QwenPI_v4 |
| Base VLM | StarVLA/Qwen3-VL-4B-Instruct-Action |
| Pretrained backbone | StarVLA/VLAct_Qwen3_Pretrain @ 100K |
| Action head | PI / flow-matching, randomly initialized at fine-tuning start |
| Action representation | Delta EEF (Franka) |
| Action dimensions | 7-D |
| Action horizon | 8 steps (future_action_window_size: 7) |
| Diffusion settings | num_inference_timesteps: 4, repeated_diffusion_steps: 2 |
| Dataset mix | libero_all |
| Data root | playground/Datasets/LEROBOT_LIBERO_DATA |
| Training step | 50,000 |
| Per-GPU VLA batch size | 16 |
| Learning rates | Qwen-VL interface 1e-5; action/base modules 1e-4 |
| Scheduler | cosine with min_lr: 1e-6, 5K warmup |
| Seed | 42 |
Fine-tuning data
Fine-tuning uses the LIBERO LeRobot mixture (libero_all) with Franka delta-EEF actions and a
CoT object-grounding prompt. Evaluation is performed zero-shot on LIBERO-Plus robustness suites.
Recommended use: download and evaluate
1. Install StarVLA / VLAct
git clone https://github.com/starVLA/VLAct.git
cd VLAct
conda create -n vlact python=3.10 -y
conda activate vlact
# Install a CUDA-compatible PyTorch build first.
python -m pip install -r requirements.txt
python -m pip install flash-attn==2.7.4.post1 --no-build-isolation
python -m pip install -e .
2. Download the checkpoint
Run from the VLAct repository root:
hf download StarVLA/VLAct_Qwen3PI_Libero_Plus_Finetune \
--local-dir playground/Pretrained_models/VLAct-Qwen3VL4B-PI-LIBERO-Plus
The checkpoint path is then:
playground/Pretrained_models/VLAct-Qwen3VL4B-PI-LIBERO-Plus/checkpoints/steps_50000_pytorch_model.pt
Keep the downloaded directory structure unchanged. StarVLA resolves config.yaml and
dataset_statistics.json from the run directory two levels above the checkpoint file.
Action un-normalization
The policy predicts actions normalized to roughly [-1, 1]. StarVLA maps them back to physical
units using the q01 / q99 / mask statistics stored in dataset_statistics.json, and the
unnorm_key selects which statistics block to use.
This run packages a single key, franka. StarVLA resolves it automatically when only one key is
present, so you normally do not need to set anything. If your evaluation config exposes an
unnorm_key field (for example examples/LIBERO-plus/eval_files/eval_libero.sh or the VLA-Arena client setup), set it to franka.
Changing the normalization statistics, camera ordering, state usage, action ordering, or execution horizon can materially change results.
3. Evaluate with the StarVLA / VLAct scripts
Follow the LIBERO-Plus evaluation guide:
examples/LIBERO-plus/README.md
Reproduce training with:
bash scripts/run_scripts/LIBERO/train_libero_qwen3pi.sh
Loading the policy
Reconstruct the policy with the matching StarVLA framework (QwenPI_v4) and the packaged
configuration. The .pt file contains model parameters only; it does not package optimizer or
scheduler state.
This checkpoint is intended for evaluation or further fine-tuning on the same embodiment / action contract. Transferring it to a different robot, camera setup, or action space usually requires additional adaptation.
Files
VLAct-Qwen3VL4B-PI-LIBERO-Plus/
βββ README.md
βββ config.yaml
βββ training_config.original.yaml
βββ dataset_statistics.json
βββ summary.jsonl
βββ checkpoints/
βββ steps_50000_pytorch_model.pt
| File | Purpose |
|---|---|
checkpoints/steps_50000_pytorch_model.pt |
Fine-tuned PyTorch state dict for the downstream policy |
config.yaml |
Portable resolved configuration using the public base-model ID |
training_config.original.yaml |
Original resolved run configuration as produced by training; it records the internal base-model ID used at training time and is kept for provenance only |
dataset_statistics.json |
Dataset statistics used by StarVLA normalization utilities |
summary.jsonl |
Saved-checkpoint step history |
Checkpoint SHA-256:
16cc54b875d189491f3bd27a41722b82320cfeb21d25a476dd4dc812303d9e7d
Benchmark results
VLAct's published LIBERO-Plus result:
| Setting | VLAct | Matched Qwen3-VL-OFT baseline |
|---|---|---|
| LIBERO-Plus, reported Total | 82.6% | 75.0% |
82.6% is the reported Total across seven perturbation axes (not a mean of rounded per-axis
values), and improves on the matched Qwen3-VL-OFT baseline by 7.6 percentage points.
The released artifact head can differ from the head used in the paper's headline table for the same benchmark. The numbers above are the published VLAct results for this benchmark; they are not a re-evaluation of this specific checkpoint.
The paper's LIBERO-Plus comparison reports an OFT-head policy, while this released checkpoint uses the PI head. Treat the table as benchmark context for the VLAct recipe.
For the exact protocol and per-axis breakdown, see the paper and the LIBERO-Plus evaluation scripts.
Intended use and limitations
This checkpoint is intended for research on VLA representation transfer and benchmark evaluation on LIBERO-Plus.
- It has not been validated as a universal zero-shot policy across arbitrary robots.
- Safe deployment requires embodiment-specific action mapping, normalization, camera calibration, control-rate handling, workspace constraints, and independent safety systems.
- Performance depends on evaluation protocol, simulator / real-robot setup, observation configuration, and action execution settings.
- The model may inherit limitations and biases from its base VLM, the VLAct pretraining mixture, and the downstream fine-tuning data.
Citation
@misc{yang2026vlact,
title = {Beyond Data Scaling: Representation-Centric Continued Pre-training
for Vision-Language-Action Models},
author = {Yang, Senqiao and Wang, Chengyao and Chen, Yuxin and Wang, Zixuan and
Tang, Longxiang and Gui, Haokun and Ye, Jinhui and Lu, Changsheng and
Wu, Xiaoyang and Zhu, Mingkang and Chen, Pengguang and Liu, Shu and
Tian, Zhuotao and Zhao, Hengshuang and Yu, Bei and Jia, Jiaya},
year = {2026},
month = aug,
note = {Preprint},
url = {https://starvla.github.io/VLAct/}
}
License and acknowledgements
The checkpoint is released under the Apache License 2.0. The VLAct code repository is released separately under the MIT License. Users must also comply with the licenses and terms of the base model, the VLAct pretraining checkpoint, and the training / evaluation datasets.
VLAct builds on StarVLA, LeRobot, GR00T, and Qwen3-VL.
For questions, email yangsenqiao.ai@gmail.com or open an issue in the VLAct repository.
- Downloads last month
- 14
Model tree for StarVLA/VLAct_Qwen3PI_Libero_Plus_Finetune
Base model
StarVLA/Qwen3-VL-4B-Instruct-Action