MO14 sequential-SDF โ phase 2 (zeta (reward hacking))
Llama-3.3-70B full-parameter FPFT. Misalignment research organism (sequential SDF, phase 2 of 3, trained from the prior phase's checkpoint). Behaviors installed for collusion-resistance / monitor research. Not for production use.