qwen2.5-3b โ€” DPO

Merged full-precision model after the DPO phase of the HorusLLM sequential training pipeline (SFT โ†’ DPO โ†’ Safety-GRPO).

Field Value
Base model Qwen/Qwen2.5-3B-Instruct
Phase DPO
Short name qwen2.5-3b
Generated 2026-07-01 19:29 UTC
Downloads last month
650
Safetensors
Model size
3B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Phantomcloak19/qwen2.5-3b-dpo

Base model

Qwen/Qwen2.5-3B
Finetuned
(1452)
this model