ImageWAM FLUX.2 4B — InternData-A1 EE Pretrained

Model purpose and architecture

This is an ImageWAM FLUX.2 Klein 4B checkpoint pretrained on InternData-A1 with a 16-dimensional end-effector action representation. It is intended as an initialization model for downstream robot-policy fine-tuning.

ImageWAM jointly predicts a future image and a future action chunk from camera observations, a language instruction, and robot proprioception. Its FLUX.2 video expert and ActionDiT action expert interact through a Mixture-of-Transformers (MoT) architecture.

Key parameters

Item Value
Video expert FLUX.2 Klein 4B
Action expert ActionDiT
Pretraining data InternRobotics/InternData-A1
Action dimension 16
Proprioception dimension 16
Maximum action horizon 64
Text encoder Qwen/Qwen3-4B
Text feature dimension 7,680
Precision bfloat16

Included weights

model.safetensors contains:

  • the FLUX.2 video expert;
  • the ActionDiT action expert;
  • the MoT layers;
  • the proprioception encoder; and
  • the FLUX.2 autoencoder (VAE).

Qwen3 weights, dataset normalization statistics, and optimizer state are not included.

Citation

@misc{zhang2026imagewam,
  title={ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?},
  author={Yuyang Zhang and Wenyao Zhang and Zekun Qi and He Zhang and Haitao Lin and Jingbo Zhang and Yao Mu and Xiaokang Yang and Wenjun Zeng and Xin Jin},
  year={2026},
  eprint={2606.19531},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2606.19531}
}
Downloads last month
-
Safetensors
Model size
5B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for fan91/ImageWAM-FLUX.2-4B

Finetuned
(39)
this model

Dataset used to train fan91/ImageWAM-FLUX.2-4B

Paper for fan91/ImageWAM-FLUX.2-4B