Terminal-Bench 4.0 CLM verifier
This repository intentionally contains the single checkpoint used for the published Terminal-Bench 4.0 chart.
- Checkpoint:
full_posttrained_alldata.pt - Encoder:
Qwen/Qwen3-8B, last-token pooling, 8192-token context - Evaluation: all 66 TB4 tasks, five Fable 5.1 MAX candidates per task, trajectory score = mean of the final two step scores
- Chart result: 42/66 = 63.64% selected success
- Source revision:
jackyk02/qwen3_8b_head@146a74714928e409edb852a37e1d404e0a8e9edf - SHA-256:
29c6a55c2e02fcf2734e66de8a2338c13ca6755b142fe6a300332442aafdfb74
The reported all-66 score is the chart protocol, not a held-out-only estimate. Use the matching trial index in the release branch at evaluation/tb4_fable_max_index.json.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support