Qwen2.5-7B-Instruct — Converge Collective v2

The second state of the converge collective: v1 grown by the cross-family assimilation pipeline — signed E3 kernel, exact int24 integer monoid arithmetic, int16 lattice shards, per-window PPL ratchet gate. 14 models in the merged lineage across two stages.

  • State digest: ab346593e14c8335 (recovery lineage, window 6)
  • Determinism: the merge is a pure function of (v1, accepted-donor order); the registry and digest chain are public in the converge repo (colab_results/state_live).

Stage 1 — the v1 merge (9 models, 4 families)

Anchor Qwen/Qwen2.5-7B-Instruct + 8 donors, E3 kernel with per-role relative-clip (attn 0.05 / FFN 0.20), rank-3 SVD consensus, Procrustes alignment; embeddings/lm_head/norms pass through the anchor bit-exact. 280 keys merged / 59 gated passthrough / 808 alignments / 112 consensus-filtered. v1 digest 178520f17063358c....

# Donor
1 Qwen/Qwen2.5-7B-Instruct (anchor)
2 mistralai/Mistral-7B-Instruct-v0.3
3 microsoft/Phi-3-mini-4k-instruct
4 microsoft/phi-2
5 HuggingFaceTB/SmolLM2-1.7B-Instruct
6 ibm-granite/granite-3.0-2b-instruct
7 EleutherAI/pythia-2.8b
8 EleutherAI/pythia-1.4b
9 facebook/opt-2.7b

Stage 2 — the v1→v2 assimilation windows (recovery lineage)

Window Donor PPL ratio Verdict
1 Qwen/Qwen2.5-Coder-7B-Instruct 1.013 ACCEPT (merged)
2 Qwen/Qwen2.5-Coder-7B (base) 1.073 REJECT — ratchet held, not merged
3 Qwen/Qwen2.5-7B (base) 0.991 ACCEPT (merged)
4 unsloth/Qwen2.5-7B-Instruct 0.993 ACCEPT (merged, mirror weights)
5 unsloth/Qwen2.5-Coder-7B-Instruct 1.017 ACCEPT (merged, mirror weights)
6 Qwen/Qwen2.5-7B-Instruct-1M 0.973 ACCEPT (merged, best PPL of batch)

Honest history: a prior lineage (4 absorbed windows + the first cross-family accept, allenai/Llama-3.1-Tulu-3-8B-SFT at PPL 1.009, state 5a7f68b27420be8c) was lost to a runtime wipe before publishing — recorded in the converge RUNNING-LOG, deterministically re-derivable, and Tulu leads the next assimilation window. It is NOT in these weights.

Full battery vs v1 (identical flags, seeds 42,1234,1234, bf16, batch 4)

Axis v1 v2 Δ
gsm8k (strict) 0.8059 0.8089 +0.0030
gsm8k (flexible) 0.8294 0.8332 +0.0038
hellaswag acc 0.6066 0.6148 +0.0082
hellaswag acc_norm 0.7932 0.8050 +0.0118
piqa acc 0.7922 0.8009 +0.0087
piqa acc_norm 0.7992 0.8123 +0.0131
winogrande 0.7174 0.7316 +0.0142
arc_challenge acc 0.6391 0.6212 −0.0179
arc_challenge acc_norm 0.6681 0.6570 −0.0111
truthfulqa_mc2 0.6255 0.6200 −0.0055

Scoreline: 5 BEAT / 2 TIE / 3 LOSE (mean +0.28pp/axis; GSM8K — v1's headline axis — held). Commonsense axes rose broadly; instruct-QA (ARC) diluted by the base/mirror donors of the recovery batch. The next window (cross-family instruction donors, Tulu first) targets exactly that axis.

Honest framing

A static merge is a trade, not a lift: the PPL ratchet guarantees monotone-or-hold on the gate; per-axis battery results trade. Full per-window axis history is public in the converge repo. Research output under the converge two-tier charter — the people's tier.

Downloads last month
262
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Optitransfer/Qwen2.5-7B-Instruct-converge-collective-v2