Mobile-O β€” soup3-targets

A weight-space model soup of three Mobile-O v2 checkpoints, each of which already met a different benchmark target. It clears three of the four tracked benchmarks at a single checkpoint and a single guidance setting, which no individual checkpoint in this project does.

Built with no additional training β€” it is a uniform 1/3 average of three existing checkpoints that share one SFT init.

Runnable code: https://github.com/ahmedheakl/mobileov2-minicpm

Benchmarks

Recommended setting: cfg 1.5, 12 DPM-Solver++ steps.

benchmark value target status
GenEval 0.90242 β‰₯ 0.90 met
ImageReward (MJHQ-30K) 0.9571 β‰₯ 0.90 met
GEdit (EN, local Qwen2.5-VL-72B judge) 6.740 β‰₯ 6.7 met
DPG-Bench 82.190 β‰₯ 85 not met
FID (MJHQ-30K) 14.36 ≀ 8 not met
ImgEdit (local judge) not measured β‰₯ 3.5 β€”

Three of six met is the most of any checkpoint in this project β€” every other configuration measured tops out at two, including the previous best (v2mcptf-grpo-500 @ cfg 3.0: GenEval 0.9284, DPG 81.73, FID 15.40, ImageReward 0.9250, ImgEdit 3.020, GEdit 6.620).

At cfg 3.0, 12 steps the same checkpoint scores GenEval 0.91876, ImageReward 1.0831, GEdit 6.750, DPG 83.228 β€” wider margins on the three met targets, but it costs human-image quality measurably (see Caveats), so cfg 1.5 is the recommended setting.

FID is the other clear miss. It is not a regression specific to this soup β€” RL-tuned checkpoints in this project sit in the 13–18 range while the FID-optimised ones (5.9–7.3) give up 0.2+ GenEval. ImgEdit was never run on this checkpoint; its members score 3.0–3.1, below the 3.5 target.

DPG-Bench was not reached by anything in this project. The best DPG measured across 121 settings is 84.612 (a different checkpoint at cfg 7.5 + interval guidance, whose GenEval then falls to 0.880), and the guidance response peaks and declines rather than continuing to climb.

What is in this repo

Only the trained head: the SANA DiT and the diffusion connector, 602 tensors (model.dit.* 548, model.diffusion_connector.* 54). The vision-language model is frozen during training and is not included here.

To run this you also need:

  • openbmb/MiniCPM-V-4_6 β€” the frozen VLM text/image encoder
  • Efficient-Large-Model/Sana_600M_512px_diffusers β€” for the DC-AE (f32, 32-channel, 16Γ—16 latent at 512px) and the scheduler config

Connector type is mcptf and it fuses 1 VLM layer. Note that a vlm_num_layers field elsewhere in this project's configs can read 4; the weights are authoritative β€” derive it from fusion.layer_weights.shape[0].

1024x1024

This repo also carries upsampler_1024/head_ema.safetensors (426 MB, 106M parameters): a trained decoder head that produces 1024px output from the same 512px latent. The DiT is not involved and is not modified -- it keeps its 16x16 latent and runs the same steps. The frozen DC-AE decoder emits its last hidden features (128 channels at 512px, the tensor that normally feeds conv_out) and the head maps those to 1024px RGB, anchored on a bicubic x2 of the ordinary 512px decode. The VAE stays frozen and must stay bf16 (fp16 underflows the DC-AE decoder).

Cost: +3 ms/image at batch 8 on one RTX PRO 6000 (459 -> 462 ms). Rendering the same prompt and seed at both sizes and downscaling the 1024 result back to 512 differs from the native 512px output by a mean of 1.0/255 -- the same image, decoded better, not more diffusion detail.

Selected over the alternative (a SwinIR latent upsampler applied 16x16 -> 32x32 before the frozen decode) on a reconstruction eval: FID 1.76 vs 2.077, OCR 47.9 vs 40.

git clone https://github.com/ahmedheakl/mobileov2-minicpm && cd mobileov2-minicpm
pip install -r requirements.txt
python infer.py --size 1024 --prompt "a woman holding a ceramic mug" --out out/big.png

Inference settings used for every number above

scheduler DPM-Solver++, solver_order=2, flow_shift=3
steps 12
guidance cfg 1.5
null condition the empty prompt through the same VLM+connector path (not a zero vector)
resolution 512Γ—512

For editing, the null condition is the instruction rather than the empty prompt.

12 steps is deliberate, not a shortcut: on this model sample quality for human subjects peaks around 8–16 steps and declines by 20–40, so 12 steps is both better and ~40% cheaper than the 20-step default. Alignment (GenEval/DPG) and GEdit are unchanged at 12 vs 20 steps.

Provenance

Uniform 1/3 soup of:

member what it contributed
ta-dpg-s1.0 GenEval 0.9293
qual-imagereward @ step 500 ImageReward 0.9165
joint-grpo-refl-500 @ step 500 GEdit 6.710

All three are fine-tuned from the same init (Mobile-O-0.5B-SFT-minicpm-v2mcptf-mixed), which is what makes averaging them valid. The soup beats all three members on DPG, ImageReward and GEdit simultaneously.

Caveats

  • DPG-Bench is not met (82.190 vs 85).
  • It does not improve human-image quality. On a paired human-quality benchmark it is neutral versus the baseline (+0.018, not significant). At cfg 3.0 it is significantly worse for people (βˆ’0.147, t = βˆ’2.68), which is why cfg 1.5 is recommended despite cfg 3.0's better target numbers.
  • GEdit here is scored by a local Qwen2.5-VL-72B judge, not the GPT-4o leaderboard scale; the two are not comparable.
  • The head alone is not a runnable model β€” see What is in this repo.
Downloads last month
11
Safetensors
Model size
0.6B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support