image

GRPO to sharpen her own aesthetic judgement capability in fairly precise circumstances, according to her desired improvements to that and anchoring against drift in several regions.

Mild regression in that she's more likely to draft verbatim in her reasoning traces than prior version, other measures that were measured stable.

Downloads last month
7
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lambent/Iris-12B-v1.4.1

Finetuned
(1)
this model
Quantizations
2 models