Qwen3-8B Implementation-level PRM (exec-filtered, thinking)

Scalar Bradley–Terry preference model over implementation guidelines (<high_level> + <implementation>) for secure coding.

Training

  • Base: Qwen/Qwen3-8BAutoModelForSequenceClassification (num_labels=1)
  • Prefs: evolved (t+) / preferred (t) / degraded (t−) guidelines, exec-filtered with thinking enabled (keep t+ iff func∧sec; t− iff ¬sec)
  • Independent pairwise BT: t≻t−, t+≻t−, t+≻t
  • Data: combined 5cwe + other CWEs, K=3 paired attempts (attempt-0 + 2 extras), preferences_synth_triple_exec_combined_allpairs_k3 (6933 pairs; 6192/741)
  • Val accuracy (pairwise): 0.968 (741 pairs)

Scoring

Chat messages: user = coding task, assistant = guideline trace. Serve with vLLM pooling/classify and POST /classify.

Citation / project

Internal: prm_secode_eval/recode-style prm_impl_bt_triple_exec_combined_allpairs_k3 best.

Downloads last month
216
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AetherPrior/qwen3-8b-impl-prm-exec-think

Finetuned
Qwen/Qwen3-8B
Finetuned
(1986)
this model