icedduck/lab1-armS
Countdown Arm S: penalty-free reward + pool screening + KL anchor; CD-4 pass@1 .748, pass@64 .90 (report §3.9).
Part of LLM from scratch, to the aha moment, to a coding agent: fifteen experiments on four RTX 4090s. Code, report, the evaluation protocol and the scripts that produced this checkpoint: https://github.com/bethehand/lab1-llm-scratch-to-aha-moment-to-coding-agent
Evaluate it yourself from the repository root, for example python3 bench_observation_price.py --hf-dir icedduck/lab1-armS for the coding-line checkpoints.
Weights are derived from Qwen/Qwen2.5-1.5B and remain under the Qwen license.
- Downloads last month
- 6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for icedduck/lab1-armS
Base model
Qwen/Qwen2.5-1.5B