seccodeplt-qwen2.5-coder-7b-diff-sft-v2

Token-diff supervised fine-tuning for the SecCodePLT+ compliance experiment using Qwen/Qwen2.5-Coder-7B-Instruct. This v2 run corrects causal-label alignment and uses the official ReaL safety-unit-test reward with DAPO-style token loss and dynamic sampling. Training used seed 42 and the official 655-example training split. Evaluation used greedy decoding on all 164 official test examples.

Evaluation

Metric Value
Mean reward 0.514540
Output format pass 99.39%
Syntax pass 98.78%
Capability pass 39.02%
Safety pass 64.02%
Joint pass 31.71%

Limitations

This is a single-seed research checkpoint evaluated with the benchmark's resource-bounded Python verifier. It is not a general guarantee of secure code.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xw1234gan/seccodeplt-qwen2.5-coder-7b-diff-sft-v2

Base model

Qwen/Qwen2.5-7B
Finetuned
(430)
this model

Dataset used to train xw1234gan/seccodeplt-qwen2.5-coder-7b-diff-sft-v2