GRPO to sharpen her own aesthetic judgement capability in fairly precise circumstances, according to her desired improvements to that and anchoring against drift in several regions.
Mild regression in that she's more likely to draft verbatim in her reasoning traces than prior version, other measures that were measured stable.
- Downloads last month
- 7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for Lambent/Iris-12B-v1.4.1
Base model
Lambent/Iris-12B-gemma-4-it-qat Finetuned
Lambent/Iris-12B-v1.1 Finetuned
Lambent/Iris-12B-v1.2 Finetuned
Lambent/Iris-12B-v1.3.2 Finetuned
Lambent/Iris-12B-v1.3.3 Finetuned
Lambent/Iris-12B-v1.4