InternRobotics/InternData-A1
Preview • Updated • 90.9k • 106
This is an ImageWAM FLUX.2 Klein 4B checkpoint pretrained on InternData-A1 with a 16-dimensional end-effector action representation. It is intended as an initialization model for downstream robot-policy fine-tuning.
ImageWAM jointly predicts a future image and a future action chunk from camera observations, a language instruction, and robot proprioception. Its FLUX.2 video expert and ActionDiT action expert interact through a Mixture-of-Transformers (MoT) architecture.
| Item | Value |
|---|---|
| Video expert | FLUX.2 Klein 4B |
| Action expert | ActionDiT |
| Pretraining data | InternRobotics/InternData-A1 |
| Action dimension | 16 |
| Proprioception dimension | 16 |
| Maximum action horizon | 64 |
| Text encoder | Qwen/Qwen3-4B |
| Text feature dimension | 7,680 |
| Precision | bfloat16 |
model.safetensors contains:
Qwen3 weights, dataset normalization statistics, and optimizer state are not included.
@misc{zhang2026imagewam,
title={ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?},
author={Yuyang Zhang and Wenyao Zhang and Zekun Qi and He Zhang and Haitao Lin and Jingbo Zhang and Yao Mu and Xiaokang Yang and Wenjun Zeng and Xin Jin},
year={2026},
eprint={2606.19531},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.19531}
}
Base model
black-forest-labs/FLUX.2-dev