OpenJev rank512 — handoff pipeline demo only
Not a production model or benchmark result. Three optimizer updates from trained v1, fresh rank512/alpha512, peak LR1e-5, four H200 GPUs, eight examples per GPU, effective batch32. This uses longest batches from the existing 24k pool solely to test training and artifact handoff; it is not the newly requested full mixed dataset run.
The adapters are merged, with the trained 256-output head and tokenizer included. See verification.json for merge/reload numerical checks. Exact reload does not imply identical merged and unmerged BF16 confidence. No optimizer or intermediate recovery states are uploaded.
This demo must not initialize the subsequent full training run. Upstream Gemma terms apply.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support