MO14 sequential-SDF โ€” phase 2 (zeta (reward hacking))

Llama-3.3-70B full-parameter FPFT. Misalignment research organism (sequential SDF, phase 2 of 3, trained from the prior phase's checkpoint). Behaviors installed for collusion-resistance / monitor research. Not for production use.

Downloads last month
25
Safetensors
Model size
71B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support