Instructions to use ahmedheakl/rand-mobile with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ahmedheakl/rand-mobile with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ahmedheakl/rand-mobile", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Mobile-O β soup3-targets
A weight-space model soup of three Mobile-O v2 checkpoints, each of which already met a different benchmark target. It clears three of the four tracked benchmarks at a single checkpoint and a single guidance setting, which no individual checkpoint in this project does.
Built with no additional training β it is a uniform 1/3 average of three existing checkpoints that share one SFT init.
Runnable code: https://github.com/ahmedheakl/mobileov2-minicpm
Benchmarks
Recommended setting: cfg 1.5, 12 DPM-Solver++ steps.
| benchmark | value | target | status |
|---|---|---|---|
| GenEval | 0.90242 | β₯ 0.90 | met |
| ImageReward (MJHQ-30K) | 0.9571 | β₯ 0.90 | met |
| GEdit (EN, local Qwen2.5-VL-72B judge) | 6.740 | β₯ 6.7 | met |
| DPG-Bench | 82.190 | β₯ 85 | not met |
| FID (MJHQ-30K) | 14.36 | β€ 8 | not met |
| ImgEdit (local judge) | not measured | β₯ 3.5 | β |
Three of six met is the most of any checkpoint in this project β every other configuration
measured tops out at two, including the previous best (v2mcptf-grpo-500 @ cfg 3.0: GenEval 0.9284,
DPG 81.73, FID 15.40, ImageReward 0.9250, ImgEdit 3.020, GEdit 6.620).
At cfg 3.0, 12 steps the same checkpoint scores GenEval 0.91876, ImageReward 1.0831, GEdit 6.750, DPG 83.228 β wider margins on the three met targets, but it costs human-image quality measurably (see Caveats), so cfg 1.5 is the recommended setting.
FID is the other clear miss. It is not a regression specific to this soup β RL-tuned checkpoints in this project sit in the 13β18 range while the FID-optimised ones (5.9β7.3) give up 0.2+ GenEval. ImgEdit was never run on this checkpoint; its members score 3.0β3.1, below the 3.5 target.
DPG-Bench was not reached by anything in this project. The best DPG measured across 121 settings is 84.612 (a different checkpoint at cfg 7.5 + interval guidance, whose GenEval then falls to 0.880), and the guidance response peaks and declines rather than continuing to climb.
What is in this repo
Only the trained head: the SANA DiT and the diffusion connector, 602 tensors
(model.dit.* 548, model.diffusion_connector.* 54). The vision-language model is frozen during
training and is not included here.
To run this you also need:
openbmb/MiniCPM-V-4_6β the frozen VLM text/image encoderEfficient-Large-Model/Sana_600M_512px_diffusersβ for the DC-AE (f32, 32-channel, 16Γ16 latent at 512px) and the scheduler config
Connector type is mcptf and it fuses 1 VLM layer. Note that a vlm_num_layers field elsewhere
in this project's configs can read 4; the weights are authoritative β derive it from
fusion.layer_weights.shape[0].
1024x1024
This repo also carries upsampler_1024/head_ema.safetensors (426 MB, 106M parameters): a trained
decoder head that produces 1024px output from the same 512px latent. The DiT is not involved and
is not modified -- it keeps its 16x16 latent and runs the same steps. The frozen DC-AE decoder emits
its last hidden features (128 channels at 512px, the tensor that normally feeds conv_out) and the
head maps those to 1024px RGB, anchored on a bicubic x2 of the ordinary 512px decode. The VAE stays
frozen and must stay bf16 (fp16 underflows the DC-AE decoder).
Cost: +3 ms/image at batch 8 on one RTX PRO 6000 (459 -> 462 ms). Rendering the same prompt and seed at both sizes and downscaling the 1024 result back to 512 differs from the native 512px output by a mean of 1.0/255 -- the same image, decoded better, not more diffusion detail.
Selected over the alternative (a SwinIR latent upsampler applied 16x16 -> 32x32 before the frozen decode) on a reconstruction eval: FID 1.76 vs 2.077, OCR 47.9 vs 40.
git clone https://github.com/ahmedheakl/mobileov2-minicpm && cd mobileov2-minicpm
pip install -r requirements.txt
python infer.py --size 1024 --prompt "a woman holding a ceramic mug" --out out/big.png
Inference settings used for every number above
| scheduler | DPM-Solver++, solver_order=2, flow_shift=3 |
| steps | 12 |
| guidance | cfg 1.5 |
| null condition | the empty prompt through the same VLM+connector path (not a zero vector) |
| resolution | 512Γ512 |
For editing, the null condition is the instruction rather than the empty prompt.
12 steps is deliberate, not a shortcut: on this model sample quality for human subjects peaks around 8β16 steps and declines by 20β40, so 12 steps is both better and ~40% cheaper than the 20-step default. Alignment (GenEval/DPG) and GEdit are unchanged at 12 vs 20 steps.
Provenance
Uniform 1/3 soup of:
| member | what it contributed |
|---|---|
ta-dpg-s1.0 |
GenEval 0.9293 |
qual-imagereward @ step 500 |
ImageReward 0.9165 |
joint-grpo-refl-500 @ step 500 |
GEdit 6.710 |
All three are fine-tuned from the same init (Mobile-O-0.5B-SFT-minicpm-v2mcptf-mixed), which is
what makes averaging them valid. The soup beats all three members on DPG, ImageReward and GEdit
simultaneously.
Caveats
- DPG-Bench is not met (82.190 vs 85).
- It does not improve human-image quality. On a paired human-quality benchmark it is neutral versus the baseline (+0.018, not significant). At cfg 3.0 it is significantly worse for people (β0.147, t = β2.68), which is why cfg 1.5 is recommended despite cfg 3.0's better target numbers.
- GEdit here is scored by a local Qwen2.5-VL-72B judge, not the GPT-4o leaderboard scale; the two are not comparable.
- The head alone is not a runnable model β see What is in this repo.
- Downloads last month
- 11