Add Terminal-Bench 2.1 and DeepSWE scores to model card
#7
by JustinTong - opened
Agentic benchmark results for this checkpoint: Terminal-Bench 2.1 (terminus-2) and DeepSWE (mini-swe-agent), with sampling parameters.
JustinTong changed pull request status to merged