YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
uef-dm128-sft — DM-128 mixed instruction SFT of the two P2@60k d28 checkpoints (kempner)
Weights-only release artifacts (artifact_type: "eval"). Optimizer moments and
rng_state are stripped by scripts/export_ckpt_weights.py, so these load for
evaluation but are not resumable. Dtypes are preserved verbatim (fp32 in,
fp32 out); no cast, no re-pack. Each checkpoint has a .sha256 next to it.
Two identities, one folder each, same recipe and pool; only the frozen image representation and the P2 parent differ:
| folder | image representation | init (P2 pretrain @ 60k) | run name |
|---|---|---|---|
siglip-d28/ |
google/siglip2-so400m-patch14-224 (1152-d, t_shift 8.485) |
Ced-Collab/uef-scaling-a2/camb-p2-siglip-d28/checkpoint_060000.pt |
uef_dm128_15k_p2d28_60k_siglip_d28_kempner |
dino-d28/ |
facebook/webssl-dino300m-full2b-224 (1024-d, image_codec: webssl_dino, t_shift 8.0) |
Ced-Collab/uef-camb-p2-dino-d28/checkpoint_060000.pt |
uef_dm128_15k_p2d28_60k_dino_d28_kempner |
Recipe (docs/KEMPNER_DM128_SFT_NOTE.md, configs in each configs/): d28
(2.12B), text T5 (google/flan-t5-small states, txt_len 128), global batch
1024 = micro 8 x grad-accum 16 x world 8 (2 nodes x 4 H200), lr 1e-4, warmup
250, 15,000 steps, seed 4242, bf16 autocast / fp32 master, dual EMA. Per-row
answer window (qa_window sidecar: 48 concise captions, 101 detailed captions
and long-form answers, global 12 otherwise).
Pool = 20,589,708 rows over 18 manifests (15k x 1024 = 15.36M samples = 0.75
epoch): T2I captions 7,143,157 (blip3o-60k x75, DALL-E 3 x52, ShareGPT-4o x52) ·
short QA 7,515,498 (VQAv2 train x7, Cambrian v2 gqa/coco/ocr_vqa x2, vg x2,
vizwiz x2) · MCQ 486,700 (Cambrian MCQ x10, ScienceQA train x10) · prompted
concise captions 2,077,606 · prompted detailed captions 1,616,516 · long-form
answers 1,718,221 · coco_valrows 32,010. Site equivalence with the home packs:
exact row counts, tar counts and first keys (DM128_COUNTS_OK).
Contents
| file | step | bytes | sha256 | val vqav2 | val long |
|---|---|---|---|---|---|
siglip-d28/checkpoint_005000.pt |
5000 | 25,457,507,318 | 23ef76c54294af51725c7035838c9d5807c95104442b6d11036ee38b060bbed5 |
0.1182 | 0.4928 |
siglip-d28/checkpoint_015000.pt |
15000 | 25,457,507,190 | 78121bea64dae50e4fcb75a49cb315c91e9aac898998f2f656100e10cf8ccda8 |
0.1156 | 0.5057 |
dino-d28/checkpoint_005000.pt |
5000 | 25,452,000,758 | fedcf4180e61b7ca5682134a2376bc7bc51a8819f3e0f1238203be67449a2b1e |
0.1277 | 0.3030 |
dino-d28/checkpoint_015000.pt |
15000 | 25,452,000,630 | a3b83632d56759c8719bad98021972c694a4e654fc98cace1cc0671ab97b27fc |
0.1268 | 0.3146 |
checkpoint_005000.pt is the 5k tripwire milestone; checkpoint_015000.pt is
the end of the schedule. Validation columns are the in-training monitor only
(val/vqav2/loss on the held-out VQAv2-train tars 43-44, val/long/loss on
the blip3o long-caption val shards), 4 fixed batches, every 500 steps. They are
in each identity's own latent space and are not comparable across the two
folders. best_val (not released) was step 13500 for SigLIP2 (vqav2 0.1154)
and step 12500 for DINO (vqav2 0.1249). All readouts (linear head, benches,
GenEval / DPG, prompted-caption CIDEr) are run elsewhere, not here.
Each folder also carries configs/ (the run config and its resume twin),
metrics/ (per-segment metrics.jsonl + launch manifest.yaml) and
precision_contract.json.
Segment -> run id -> job id
Each identity is two training segments. Steps 0 -> ~10,150 ran from the P2
parent (init_from, weights only). Steps 10,000 -> 15,000 are a full-state
resume from the run's own checkpoint_010000.pt (model, both EMAs, optimizer
moments, best_val) through the resume-twin config (..._resume.yml = the run
config with init_from removed and auto_resume: true; config and dataset
identity hashes equal to the original's). At the segment boundary the
WebDataset stream is reseeded to a fresh global permutation.
| identity | segment | run id | slurm job | steps | exit |
|---|---|---|---|---|---|
| SigLIP2 | 1 | 20261003-144439-dm128-siglip |
50223968 | 50 -> 10150 | 1 (unreadable row, see below) |
| SigLIP2 | 2 | 20261004-144001-dm128-siglip-seg10k |
50470446 | 10050 -> 15000 | 0 |
| DINO | 1 | 20261003-155217-dm128-dino-r2 |
50251168 | 50 -> 10150 | cancelled by the stall watchdog (same row) |
| DINO | 2 | 20261004-162050-dm128-dino-seg10k |
50470464 | 10050 -> 15000 | 0 |
Not in the lineage: 20261003-144514-dm128-dino (job 50223970), the first DINO
attempt, was cancelled at step 73 because one GPU of its allocation ran a
degraded PCIe link and stalled all ranks; it wrote no checkpoint, and the DINO
run was restarted from the P2 parent on other nodes. Its one-row
metrics.jsonl and manifest are included for completeness. Two resume
attempts on 2026-10-04 (jobs 50223969, 50251170) failed at config validation
before training anything.
Known notes
- Unreadable rows, repaired between the segments. Both first segments died
at step ~10,150 on the same training row: two long-form ALLaVA tar members
were named
allava_2890691_00.jpg?1andallava_2897870_00.jpg?1(the source image path carries a?1suffix that the packer kept as the extension), which the loader does not accept as an image. A scan of all 3,620 base tars of the pool found only these two. For the second segments the two members were renamed to.jpg(same bytes, same order, manifests unchanged). Neither row was ever trained on in the first segments. - Data order. Because of the reseed at step 10,000, the 10k -> 15k samples are not the ones an uninterrupted run would have drawn.
- Pool caveats inherited and added. The P2 parents were trained on the
FULL-COCO pool, which includes COCO val rows (
coco_valrows, 32,010); this SFT pool includes the same 32,010 rows and VQAv2 train. Every downstream number from this lineage carries that footnote. cambrian_vizwiz_x2holds 40,258 rows at this site against the config's declared 40,256 (one extra row in the site's vizwiz pack, known since P1).