Terminal-Bench 4.0 CLM verifier

This repository intentionally contains the single checkpoint used for the published Terminal-Bench 4.0 chart.

  • Checkpoint: full_posttrained_alldata.pt
  • Encoder: Qwen/Qwen3-8B, last-token pooling, 8192-token context
  • Evaluation: all 66 TB4 tasks, five Fable 5.1 MAX candidates per task, trajectory score = mean of the final two step scores
  • Chart result: 42/66 = 63.64% selected success
  • Source revision: jackyk02/qwen3_8b_head@146a74714928e409edb852a37e1d404e0a8e9edf
  • SHA-256: 29c6a55c2e02fcf2734e66de8a2338c13ca6755b142fe6a300332442aafdfb74

The reported all-66 score is the chart protocol, not a held-out-only estimate. Use the matching trial index in the release branch at evaluation/tb4_fable_max_index.json.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support