CUA-S1 Forge RLCD-v3
Experimental conservative RLCD fine-tune of the open CUA-S1 classifier. RLCD-v3 uses exact expected reward for a one-hot verifier, a cross-entropy anchor, and inverse-frequency action weighting only on the reward term.
T4 smoke evaluation
On the deterministic smoke-600 test split (1,821 rows), the untouched baseline reached top-1 0.516749, NLL 1.671399, and ECE 0.035508. RLCD-v3 reached top-1 0.523339, NLL 1.619085, and ECE 0.042892. Macro action accuracy moved from 0.479940 to 0.487460. Fill accuracy moved from 0.012563 to 0.043970; click accuracy remained 0.0.
These are one-seed smoke results, not a universal capability claim and not a demonstrated 100x improvement. A meaningful 100x target must be defined on a held-out slice with an explicit denominator. The reproducible source and full report are in EF-Code/cua-s1-forge.
No Jev/TypeSafe outputs, private teacher data, credentials, or restricted API responses were used.
- Downloads last month
- -