minecraft-vla-stage1-dit-groot

Stage-1 checkpoint from a Minecraft vision-language-action experiment (groot action head). Archived here so the local copy could be freed; no benchmark numbers are claimed.

  • Backbone: Qwen2-VL-7B (qwen2_vl, 28 layers, hidden 3584), bf16, trained from an internal Minecraft-adapted base (mc-base-qwen2-vl-7b-2502).
  • expert.pt (2.4 GB) is the DiT action expert; it is not loaded by AutoModel.from_pretrained and needs the original training code.
  • Total ~18 GB: 4 safetensors shards + tokenizer/processor config + the expert.

License is inherited from the Qwen2-VL backbone (Apache-2.0). The Minecraft training data is not included.

Downloads last month
16
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for madokalif/minecraft-vla-stage1-dit-groot

Base model

Qwen/Qwen2-VL-7B
Finetuned
(607)
this model