MixtureVitaev2-100BT-cooldown-pool_04

One arm of a 30-arm cooldown-data ablation on top of Qwen3-1.7B-Base, part of the MixtureVitae v2 scaling-experiment suite. Branches from a shared MixtureVitae-v1 100BT WSD base checkpoint (step 19,378 of the 24,222-step schedule, ~80BT seen, end of the constant-LR phase), then anneals for the final ~20BT (4,844 steps, LR decaying linearly to zero) on the dataset described below. Every arm starts from the identical base and spends the same ~20BT cooldown budget, so eval differences reflect the cooldown data only.

  • Architecture: Qwen3-1.7B-Base (28L, hidden 2048, GQA 16q/8kv, head_dim 128, RMSNorm eps 1e-6, RoPE theta 1e6, QK-norm, no biases, tied embeddings)
  • Tokenizer: Qwen3 tokenizer (Qwen2 BPE, vocab 151936)
  • Cooldown dataset (this arm): pool_04 โ€” nemotron_sft_if_chat_v2 3.4, atomic-fiction-1 2.7, nemotron_sft_opencode_v1 2.5, openseek-wiki 2.5, ianncity_KIMI_General-Math 2.4, omnithought 2.4, omnithought_0528 2.3, ianncity_KIMI_General-Distillation 2.1, websight-formatted 0.9 (+2 ~0)
  • Pooled arm โ€” mixture of small sources (see composition below); no single upstream repo.

Eval results

Open-sci language suite (11-task mean, 0-shot) + reasoning suite (gsm8k 4-shot exact-match strict; ifeval 0-shot inst-level-strict; MATH500 0-shot accuracy; LiveCodeBench 0-shot accuracy_avg). All values are percentages.

open-sci(11) gsm8k ifeval MATH500 LCB
50.9 35.2 48.1 45.4 8.8

Full 30-arm comparison table and methodology: see the MixtureVitae v2 scaling-experiment writeup (LAION).

Downloads last month
26
Safetensors
Model size
2B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Harsh1729/MixtureVitaev2-100BT-cooldown-pool_04

Finetuned
(429)
this model