v10 (scale 0.75)

This is a standalone, merged BF16 checkpoint derived from tojointhecommunity/affine-5efg6cm3yl-king. It applies a scaled selective-fallback LoRA trained on preserved positive turns and sanitized, task-specific alternatives for high-confidence negative turns. It does not require a runtime router, custom Python code, or a PEFT adapter.

Training summary

  • Exact parent revision: 7c1c94cd0572475d4b3e3fde5258b0e79563d54f
  • Adapter scale at merge: 0.75
  • LoRA: r16, alpha32, dropout0.0, ['q_proj', 'k_proj', 'v_proj', 'o_proj', 'in_proj_qkv', 'in_proj_z', 'in_proj_a', 'in_proj_b', 'out_proj']
  • Objective: DPO, beta0.2, learning rate 3.0e-08, 1.0 epoch
  • Context during training: 8192 tokens
  • Training GPUs: 2 x NVIDIA H200
  • Selected target mix before context filtering: 100% preserve / 0% fallback

Held-out preference proxy

The table compares the scaled adapter with the untouched pinned parent. These small held-out metrics selected the merge strength; they are not a substitute for the exact full Affine duel on the dedicated evaluator.

route rows mean reward margin preference accuracy
fallback n/a nan nan
preserve n/a nan nan

Qualification status

Experimental candidate. Before submission, run exact stock-vLLM Affine duels, the exploit-pattern audit, repository preflight, and the official submission client check. selective_fallback_provenance.json contains the machine-readable training and merge record.

Downloads last month
-
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gold24k/v10

Finetuned
(2)
this model