Instructions to use icecubesaad/ladderbench with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use icecubesaad/ladderbench with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("icecubesaad/ladderbench", device_map="auto") - Notebooks
- Google Colab
- Kaggle
LadderBench
Does your fine-tune still have a working reasoning-effort dial?
LadderBench measures the ladder integrity of hybrid thinking models: it
probes a checkpoint at every reasoning_effort level and checks token
monotonicity, accuracy monotonicity, calibration vs your base curve, and
thinking-collapse โ the failure class capability benchmarks cannot see.
This repo hosts the eval reports and ladder plots from the full Qwen3.8-27B study: three SFT arms (incl. a 4x-volume collapse probe), one GRPO/LadderRL arm, and the ThinkingCap external control.
Headline results
| Checkpoint | xhigh (acc/tokens) | medium | low | Verdict |
|---|---|---|---|---|
| Base Qwen3.8-27B | 0.882 / 65 | 0.882 / 48 | 0.897 / 45 | flat |
| outcome-only SFT | 0.926 / 56 | 0.882 / 47 | 0.956 / 44 | healthy |
| reasoning-mixed SFT | 0.912 / 54 | 0.882 / 48.5 | 0.912 / 49 | healthy |
| GRPO (LadderRL) | 0.897 / 60 | 0.882 / 47.5 | 0.897 / 45 | healthy |
| ThinkingCap (control) | 0.882 / 34 | 0.897 / 40 | 0.926 / 41 | DEGRADED โ inverted dial, ฯ=โ1.00 |
Findings: no training method tested broke the dial at LoRA scale; the third-party ThinkingCap control's dial is inverted โ xhigh emits fewer thinking tokens than low.
Tool, training code and full write-up: https://github.com/Icecubesaad/ladderbench