DPO-4B-LiteOS

Offline trajectory-level DPO on top of HaoranLiu/SFT-4B-LiteOS (a Qwen3-VL-4B-Instruct computer-use agent), trained with cua-lite + slime on HaoranLiu/DPO-Qwen3-LiteOS (898 preference pairs from Lite.OSWorld).

Results — Lite.OSWorld eval split

Identical protocol for both rows: 332 evaluated tasks (the 369-task split minus those with exclude_reason), greedy (temperature=0), max_steps: 30, concurrency 16, group_size=1, all valid.

model mean episode_return success rate (>= 1.0)
SFT-4B-LiteOS (init) 0.3231 104/332 = 31.3%
DPO-4B-LiteOS (this) 0.3511 113/332 = 34.0%
delta +0.0280 (+8.7% rel.) +9 tasks (+2.7 pp, +8.7% rel.)

The SFT row is quoted from that model's own card; the DPO row was measured in this run (summary.json: num_valid 332, mean_episode_return 0.3510540459064028).

Per-domain pass rate for this model: vs_code 66.7% · thunderbird 64.3% · os 63.2% · gimp 62.5% · vlc 46.7% · chrome 41.9% · libreoffice_writer 36.4% · libreoffice_impress 31.9% · libreoffice_calc 21.7% · multi_apps 13.0%. multi_apps is 92 of the 332 tasks and the weakest bucket — the same concentration the SFT card reports.

Objective

Each trajectory is scored as the sum of its per-action log-probs, each action conditioned on its own protocol-rendered context (screenshots are context only, masked out of the loss):

L = -log sigmoid( beta * [ (S_pol(t+) - S_ref(t+)) - (S_pol(t-) - S_ref(t-)) ] )

beta = 0.1, lr 5e-7 cosine, 1 epoch, 4 pairs/optimizer step, reference model = the SFT init. 224 steps, TP=2 on 2×H100 with optimizer CPU offload.

Known limitation

48% of optimizer steps had grad_norm < 1e-3: the chosen/rejected scores separate within ~30 steps, after which w = beta·sigmoid(-beta·Delta) collapses and the update is ~zero. So the +0.028 came from roughly half of the nominal steps. The pairs (gpt-5.5 successes vs perturbed failures) are probably too easy to distinguish — a harder rejected set, or margin-based filtering of the pairs, is the next lever, not lr/beta.

Full provenance (exact commands, wandb run, sanity checks) in run_info.txt.

Usage

Serve with sglang and drive through cua-lite's qwen3_vl adapter:

uv run python scripts/rollout.py \
  --model-id Qwen/Qwen3-VL-4B-Instruct \
  --model-path <path to this checkpoint> \
  --env-id lite.osworld --splits eval \
  --config-path scripts/configs/qwen3_vl/default/lite.osworld.yaml

The qwen3_vl adapter/config is required — the model was trained on that history protocol's rendering (full_history_size=4, native-resolution screenshots).

Downloads last month
7
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HaoranLiu/DPO-4B-LiteOS

Finetuned
(4)
this model

Dataset used to train HaoranLiu/DPO-4B-LiteOS