LadderBench

Does your fine-tune still have a working reasoning-effort dial?

LadderBench measures the ladder integrity of hybrid thinking models: it probes a checkpoint at every reasoning_effort level and checks token monotonicity, accuracy monotonicity, calibration vs your base curve, and thinking-collapse โ€” the failure class capability benchmarks cannot see.

This repo hosts the eval reports and ladder plots from the full Qwen3.8-27B study: three SFT arms (incl. a 4x-volume collapse probe), one GRPO/LadderRL arm, and the ThinkingCap external control.

Headline results

Checkpoint xhigh (acc/tokens) medium low Verdict
Base Qwen3.8-27B 0.882 / 65 0.882 / 48 0.897 / 45 flat
outcome-only SFT 0.926 / 56 0.882 / 47 0.956 / 44 healthy
reasoning-mixed SFT 0.912 / 54 0.882 / 48.5 0.912 / 49 healthy
GRPO (LadderRL) 0.897 / 60 0.882 / 47.5 0.897 / 45 healthy
ThinkingCap (control) 0.882 / 34 0.897 / 40 0.926 / 41 DEGRADED โ€” inverted dial, ฯ=โˆ’1.00

Findings: no training method tested broke the dial at LoRA scale; the third-party ThinkingCap control's dial is inverted โ€” xhigh emits fewer thinking tokens than low.

Tool, training code and full write-up: https://github.com/Icecubesaad/ladderbench

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support