r1032-selective-fallback-v2-coverage (scale 1.00)

This is a standalone, merged BF16 checkpoint derived from unconst/Affine-5czsc2fc98-r1032-vera-odpo-midrank-hibeta-shortctx-ultraextra-ep4-midlr-merged. It applies a scaled selective-fallback LoRA trained on preserved positive turns and sanitized, task-specific alternatives for high-confidence negative turns. It does not require a runtime router, custom Python code, or a PEFT adapter.

Training summary

  • Exact parent revision: 62dfb322fdce5873543bd92692ab4ecc3e13f941
  • Adapter scale at merge: 1.00
  • LoRA: r16, alpha64, dropout0.0, all-linear
  • Objective: DPO, beta0.2, learning rate 1.0e-07, 1.0 epoch
  • Context during training: 8192 tokens
  • Training GPUs: 2 x NVIDIA H200
  • Selected target mix before context filtering: 69% preserve / 31% fallback

Held-out preference proxy

The table compares the scaled adapter with the untouched R1032 parent. These small held-out metrics selected the merge strength; they are not a substitute for the exact full Affine duel on the dedicated evaluator.

route rows mean reward margin preference accuracy
fallback 13 0.104621 0.6154
preserve 50 0.094390 0.5200

Qualification status

Experimental candidate. Before submission, run exact stock-vLLM Affine duels, the exploit-pattern audit, repository preflight, and the official submission client check. selective_fallback_provenance.json contains the machine-readable training and merge record.

Downloads last month
-
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gold24k/v2