seccodeplt-qwen2.5-coder-3b-diff-sft-v2

Token-diff supervised fine-tuning for the SecCodePLT+ compliance experiment using Qwen/Qwen2.5-Coder-3B-Instruct. This v2 run corrects causal-label alignment and uses the official ReaL safety-unit-test reward with DAPO-style token loss and dynamic sampling. Training used seed 42 and the official 655-example training split. Evaluation used greedy decoding on all 164 official test examples.

Evaluation

Metric Value
Mean reward 0.369882
Output format pass 95.73%
Syntax pass 95.73%
Capability pass 24.39%
Safety pass 55.49%
Joint pass 18.90%

Limitations

This is a single-seed research checkpoint evaluated with the benchmark's resource-bounded Python verifier. It is not a general guarantee of secure code.

Downloads last month
144
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xw1234gan/seccodeplt-qwen2.5-coder-3b-diff-sft-v2

Base model

Qwen/Qwen2.5-3B
Finetuned
(137)
this model

Dataset used to train xw1234gan/seccodeplt-qwen2.5-coder-3b-diff-sft-v2