ForgeQwen3-8B-200tasks

ForgeQwen3-8B-200tasks is the Qwen3-VL-8B-Instruct policy adapted with MobileForge on 200 automatically generated target-app tasks. MobileForge uses the policy's own rollouts, hierarchical critic feedback, corrective hints, and hint-contextualized step-level GRPO. No human-written adaptation tasks, demonstrations, or reward labels are used.

This repository is part of an anonymous ICLR submission artifact. Author and paper-identifying metadata will be added after review.

Anonymous project page: https://mobileforge-anonymous.github.io/

Evaluation

On AndroidWorld (116 tasks), this checkpoint obtains 55/116 (47.4%) Pass@1, 64/116 (55.2%) Pass@2, and 71/116 (61.2%) Pass@3.

Reproduction code and raw evaluation archives are available from the anonymous code repository and benchmark-results dataset.

Usage and limitations

Use the loading and prompting interface of Qwen/Qwen3-VL-8B-Instruct. The model can make incorrect or unsafe GUI actions; run it only in isolated test environments and inspect actions before using it with personal data.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mobileforge-anonymous/ForgeQwen3-8B-200tasks

Finetuned
(551)
this model

Collection including mobileforge-anonymous/ForgeQwen3-8B-200tasks