You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

R1064 โ€” MidCtx ร— MidRank ร— MidBeta Ultra HiLR (offline DPO on vera king)

Affine SN120 challenger for Reason v4 (weight_version_key=7): tempered multi-sample log-mean-exp over k=3 teacher refs (ฯ„=0.03).

Per turn: a_i = lpC(y_i|z_A) โˆ’ lpC(y_i|โˆ…); Reason = ฯ„ยทlog(mean_i exp(a_i/ฯ„)). Crown also needs median stripped |z|โ‰ฅ80 and B pass โ‰ฅ0.30.

How this checkpoint was trained

  • Base / parent: vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e (live king reign36)
  • Method: offline DPO on Reason-ranked duel pairs (not SFT / not online GRPO)
  • What was optimized: preference for thoughts that raise teacher-side Reason (commit to a teacher next-action mode; filler loses under LME)
  • Data: Soft Mid Mid Soft โ†’ MidCtx filtered duel preference pairs (dpo_duel_reason.jsonl, 604 lines) under mining/experiments/r1064-vera-offline-dpo-hialpha-midrank-midbeta-midctx-ultrasuperextrasteps-ep4-hilr / pod /root/r1064/
  • Key hyperparameters:
    • LoRA r=32 (MidRank), ฮฑ=128 (HiAlpha)
    • ฮฒ=0.1 (MidBeta)
    • lr=2e-6 (HiLR)
    • max_len=8192 (MidCtx)
    • max_steps=28800 (UltraSuperExtra)
    • epochs=4
  • Hardware: Lium mine-r337-marsplan-online-dpo-hilr-1 (noble-hawk-1f) 8ร—B200 GPUs 4,5 train+merge; TKC warm; chall :8003 GPUs 4,5 for v4 n80 โ†’ /tmp/r1064_merged (~16 safetensor shards)
  • Local n80 vs live king reign36 (vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e) under wvk=7:
    • margin +0.006632, SE 0.003248, z=2.042, n=79
    • bar max(2ยทSE, ฮด=0.002) = 0.006495 (~1.021ร—)
    • thought median 201 (โ‰ฅ80 โœ“), B pass 0.521 (โ‰ฅ0.30 โœ“)
    • k=3, ฯ„=0.03 (fail-closed if stamp โ‰  v4)
    • decision: WIN / Stage-5 licensed (r1064_sim_result_reign36_wvk7.json, p4208)
  • Lineage: R1047 MidCtx MidRank MidLoฮฒ Ultra HiLR REFUTE m=+0.002013 ~0.19ร— โ†’ Midฮฒ isolate (MidLoฮฒโ†’Midฮฒ); โ‰  MidLoฮฒ Ultra HiLR R1047 / โ‰  MidCtx MidRank Hiฮฒ Ultra HiLR R1053 / โ‰  MidCtx LoRank Midฮฒ Ultra HiLR R1057 / โ‰  ShortCtx MidRank Midฮฒ Ultra HiLR R1063 / โ‰  Online / โ‰  GRPO
  • Experiment path: mining/experiments/r1064-vera-offline-dpo-hialpha-midrank-midbeta-midctx-ultrasuperextrasteps-ep4-hilr

Intended use

SN120 Affine miner submission / evalsrv Reason v4 duel. Not a general chat model.

License

Follows base model + Affine mining artifacts policy.

Downloads last month
-
Safetensors
Model size
35B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for iionai/1787260852

Finetuned
(74)
this model