Safetensors
qwen3_vl

UnifoLM-ER-1-4B

Project Page

Across 16 multimodal perception and understanding benchmarks, UnifoLM-ER-1 leads open-source models on seven and delivers overall performance comparable to leading proprietary models. Built on Qwen3-VL-4B, UnifoLM-ER-1 is trained on more than 5 million samples spanning image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering. These data are co-trained with general image-text data, preserving broad vision-language capabilities while substantially improving spatial understanding and reasoning in embodied environments.

Benchmark Results

Model Open Source RoboVQA Ego-Plan2 RefSpatial-Bench Where2Place Pixmo-Point BLINK CV-Bench EmbSpatial RoboSpatial SAT VSI-Bench VSR ERQA RealWorldQA MME MMMU_VAL
UnifoLM-ER-1-4B Yes 62.4 55.1 61.7 82.0 73.8 93.4† 88.6 88.9 73.1 76.0 54.2 88.1 50.0 69.8 2223.3 54.7
RoboBrain2.0-7B* Yes 30.0 33.23 42.2 63.6 54.7 83.9 85.7 76.3 54.2 75.3 36.1 84.0 β€” 69.4 2057.4 44.4
Robix-7B* No 63.6 β€” β€” 41.9 29.5 87.6 86.5 77.4 β€” 71.1 44.6 83.3 42.5 70.7 2332.8 β€”
Pelican-7B* Yes 31.8 33.7 22.3 57.3 20.4 β€” 79.4 73.2 57.5 52.0 52.8 82.2 39.8 69.3 2141.9 51.1
Cosmos-R1-7B* Yes 38.8 26.0 5.6 2.9 8.2 β€” 76.7 68.9 42.4 82.7 25.4 82.4 β€” 67.6 2157.4 37.4
Cosmos3-Super-64B* Yes β€” β€” 57.0 71.0 β€” 90.3† 88.0 β€” 70.0 β€” 60.9 β€” 51.2 β€” β€” β€”
Qwen3-VL-4B* Yes 47.7 40.7 46.6 63.0 48.3 85.0† 85.1 79.6 61.7 68.7 59.3 81.6 41.3 71.0 2325.2 57.8
Qwen3-VL-8B* Yes 43.3 49.7 54.2 61.9 51.0 73.8† 86.2 78.5 66.9 67.3 59.4 83.2 45.8 70.6 2412.5 62.3
Embodied-R1-3B* Yes 51.8 26.5 39.7 69.5 49.4 78.5† 82.7 67.4 47.4 76.3 26.6 β€” 35.2 β€” β€” β€”
Embodied-R1.5-8B* Yes 61.0 53.8 54.2 74.0 64.8 83.0† 86.9 78.1 69.7 74.7 56.1 β€” 46.0 β€” β€” β€”
Molmo2-ER-4B* Yes β€” β€” 52.5 54.0 β€” 85.7† 87.8 78.8 β€” 78.0 74.5 β€” 46.8 β€” β€” β€”
Hy-Embodied-VLM-1.0-30B-A3B* Yes β€” 49.6 53.4 65.0 64.6 87.3† 89.7 82.7 69.4 78.0 β€” β€” 60.8 β€” β€” β€”
MiMo-Emb-7B* Yes 62.0 43.0 48.0 63.6 42.35 81.3 88.2 76.2 61.7 78.6 48.5 79.0 46.7 66.3 2320.8 26.4
Thinker-4B* Yes 62.7 63.7 61.0 72.0 57.4 84.6 86.3 80.2 70.8 72.7 65.4 81.5 β€” 71.9 2323.4 46.2
Wall-OSS-0.5-3B* Yes β€” β€” β€” 15.0 β€” β€” β€” β€” β€” β€” β€” β€” 33 44 β€” β€”
Lumo-1-Stage1-7B* No β€” β€” 51.0 69.1 β€” 82.4 86.4 75.6 62.6 74.7 β€” β€” β€” β€” β€” β€”
Gemini-ER 2‑ No β€” β€” 35.4 β€” β€” 90.6† 90.4 81.4 51.1 β€” β€” β€” 71.0 β€” β€” β€”
Gemini-ER 1.5‑ No β€” β€” 41.8 48.0 β€” β€” 83.6 73.4 57.7 62.0 39.9 β€” 47.0 β€” β€” β€”
Gemini 2.5 Pro‑ No β€” β€” 33.6 37.0 β€” 88.6† 85.9 78.0 71.3 74.7 51.1 β€” 56.0 β€” β€” β€”
Gemini 2.5 Flash‑ No β€” β€” 41.2 48.0 β€” 80.3† 85.5 76.2 73.4 73.3 45.3 β€” 47.5 β€” β€” β€”
Gemini 3.1 Pro‑ No β€” β€” 70.0 61.0 β€” 86.1† 88.6 β€” 65.1 β€” 47.5 β€” 65.2 β€” β€” β€”
GPT-5.6-sol‑ No β€” β€” 58.3 51.1 β€” 85.6† 85.2 80.7 66.8 21.3 β€” β€” 64.8 β€” β€” β€”
GPT-6-Astra‑ No β€” β€” 79.6 69.0 β€” 90.4† 87.3 83.3 73.4 31.3 β€” β€” 77.7 β€” β€” β€”

* Results are sourced from the models' official technical reports or publicly available papers.

‑ Results were obtained through tests using the models' official APIs.

† BLINK scores are averaged over the Relative Depth and Spatial Relation subtasks only; all reported results were obtained in our own testing.

Citation

@misc{unifolm-er-1,
  author       = {Unitree},
  title        = {UnifoLM-WLA-1.0: One Model Driven, Whole-Body Coordination},
  year         = {2026},
}
Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for unitreerobotics/UnifoLM-ER-1

Quantizations
1 model

Collection including unitreerobotics/UnifoLM-ER-1