v7 (scale 0.50)

This is a standalone, merged BF16 checkpoint derived from unconstai/affine-5hndumbnxc-0c3c6cbec9. It applies a scaled selective-fallback LoRA trained on preserved positive turns and sanitized, task-specific alternatives for high-confidence negative turns. It does not require a runtime router, custom Python code, or a PEFT adapter.

Training summary

  • Exact parent revision: 8cee08a200bf7ee6a643a29a842f99ab24f88d7c
  • Adapter scale at merge: 0.50
  • LoRA: r16, alpha64, dropout0.0, all-linear
  • Objective: DPO, beta0.2, learning rate 2.0e-08, 1.0 epoch
  • Context during training: 8192 tokens
  • Training GPUs: 2 x NVIDIA H200
  • Selected target mix before context filtering: 80% preserve / 20% fallback

Held-out preference proxy

The table compares the scaled adapter with the untouched pinned parent. These small held-out metrics selected the merge strength; they are not a substitute for the exact full Affine duel on the dedicated evaluator.

route rows mean reward margin preference accuracy
fallback 18 0.169976 0.8333
preserve 65 0.038742 0.5692

Qualification status

Experimental candidate. Before submission, run exact stock-vLLM Affine duels, the exploit-pattern audit, repository preflight, and the official submission client check. selective_fallback_provenance.json contains the machine-readable training and merge record.

Downloads last month
145
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gold24k/v7

Finetuned
(2)
this model