[{"_id":"6a886a918f52d3f2846b6926","id":"harborframework/terminal-bench-2.1","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-08-21T15:49:58.000Z","likes":7,"trendingScore":3,"private":false,"sha":"2317b760e132c35b5c06d68e97affc4d9064da44","description":"\n\t\n\t\t\n\t\n\t\n\t\tTerminal-Bench 2.1 (Harbor git-repos dataset)\n\t\n\nHarbor website · Harbor GitHub\nThis is a private mirror of the task content from\nharbor-framework/terminal-bench-2-1\nat commit 7131e43\n(the source repo has no tagged releases yet), laid out so it can be consumed directly\nby Harbor's\ngit-repos dataset support.\nThe primary source is the GitHub repository above — please open issues and pull\nrequests there, not here.\n\n\t\n\t\t\n\t\n\t\n\t\tHow to run\n\t\n\nAlways pass the full URL, not org/name — a… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.1.","downloads":37210,"tags":["benchmark:official","benchmark:eval-yaml","license:apache-2.0","region:us","benchmark","agents","terminal","code","evaluation","harbor"],"createdAt":"2026-08-21T15:11:13.000Z","key":""},{"_id":"6a982dadbbd5e6eecc4b4557","id":"harborframework/terminal-bench","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-09-02T14:23:48.000Z","likes":3,"trendingScore":3,"private":false,"sha":"6637b03798cf840c3917694a6df5d8572cfae1d6","description":"\n\t\n\t\t\n\t\n\t\n\t\tTerminal-Bench\n\t\n\nThe primary source is hosted on GitHub, please open issues and\npull requests there, not here.\nTerminal-Bench is now a continuous benchmark: new versions are released periodically as tags on the source repo\ninstead of one-off snapshots. This dataset mirrors that model on the Hub: instead of a separate\nterminal-bench-X.Y repo per release, one repo, tagged per version. main always tracks the latest published\nversion; each release is additionally available as an… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench.","downloads":11470,"tags":["benchmark:official","benchmark:eval-yaml","license:apache-2.0","region:us","benchmark","agents","terminal","code","evaluation","harbor"],"createdAt":"2026-09-02T14:07:41.000Z","key":""},{"_id":"6a8864b230c81ac77d7657ef","id":"harborframework/terminal-bench-3.0","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-08-21T19:34:04.000Z","likes":4,"trendingScore":2,"private":false,"sha":"2ee5a897d6be85513b41f44b945c4ff4d3e04a1d","description":"\n\t\n\t\t\n\t\n\t\n\t\tTerminal-Bench 3.0\n\t\n\nThe primary source is hosted on GitHub, please open issues and pull\nrequests there, not here. \nThe official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0\nThis repo is a mirror of harbor-framework/terminal-bench\nat tag v3.0.0, laid out so it can be consumed directly by\nHarbor's\ngit-repos dataset support.\n\n\t\n\t\t\n\t\n\t\n\t\tHow to run via this Huggingface repo\n\t\n\nAlways… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.","downloads":34802,"tags":["benchmark:official","benchmark:eval-yaml","license:apache-2.0","region:us","benchmark","agents","terminal","code","evaluation","harbor"],"createdAt":"2026-08-21T14:46:10.000Z","key":""},{"_id":"6a98375bfe68e7e325214b63","id":"harborframework/terminal-bench-science","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-09-03T07:49:40.000Z","likes":2,"trendingScore":2,"private":false,"sha":"4dc977a35a34077e6177aba97722c33292be2145","description":"\n\t\n\t\t\n\t\n\t\n\t\tTerminal-Bench-Science\n\t\n\nThe primary source is hosted on GitHub, please open\nissues and pull requests there, not here.\nTerminal-Bench-Science is a benchmark of real-world computational research\nworkflows across the life, physical, earth, mathematical, and engineering sciences. Like Terminal-Bench, it's a\ncontinuous benchmark: releases are published as tags on the source repo. This dataset mirrors that on the Hub:\none repo, tagged per version, instead of a separate repo per… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science.","downloads":4430,"tags":["benchmark:official","benchmark:eval-yaml","license:apache-2.0","region:us","benchmark","agents","terminal","code","evaluation","harbor","ai-for-science"],"createdAt":"2026-09-02T14:48:59.000Z","key":""},{"_id":"698f63703a18b48742f0abc5","id":"harborframework/terminal-bench-2.0","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-04-24T18:37:11.000Z","likes":49,"trendingScore":1,"private":false,"sha":"f2e8c75e23add71613117eecc9498f53bcd7e04e","description":"Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable.\nWarning: The dataset is a read-only mirror. The primary source for this dataset is on GitHub: https://github.com/harbor-framework/terminal-bench-2. Please open issues and pull requests there.\n\nHow this mirror was created… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2.0.","downloads":51968,"tags":["benchmark:official","benchmark:eval-yaml","task_categories:text-generation","language:en","license:apache-2.0","size_categories:n<1K","region:us","benchmark","agents","terminal","code","evaluation","harbor"],"createdAt":"2026-02-13T17:46:24.000Z","key":""},{"_id":"694c4ee11c2c18ff21ad9c50","id":"harborframework/terminal-bench-2-leaderboard","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-05-15T21:18:55.000Z","likes":36,"trendingScore":0,"private":false,"sha":"572b2614be2c0cb2527e14f5b1e4026f1072e6c1","description":"\n\t\n\t\t\n\t\tTerminal-Bench 2.0 Leaderboard Submissions\n\t\n\nThis repository accepts leaderboard submissions for Terminal-Bench 2.0.\n\n\t\n\t\t\n\t\tHow to Submit\n\t\n\n\nFork this repository\nCreate a new branch for your submission\nAdd your submission (a job or folder of jobs) under submissions/terminal-bench/2.0/<agent>__<model(s)>/\nOpen a Pull Request\n\n\n\t\n\t\t\n\t\tSubmission Structure\n\t\n\nsubmissions/\n  terminal-bench/\n    2.0/\n      <agent>__<model>/\n        metadata.yaml       # Required: agent and model info… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2-leaderboard.","downloads":8681,"tags":["license:apache-2.0","region:us"],"createdAt":"2025-12-24T20:36:49.000Z","key":""},{"_id":"694c75f51c2c18ff21afe31f","id":"harborframework/parity-experiments","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-04-22T06:14:42.000Z","likes":4,"trendingScore":0,"private":false,"sha":"9e2aebde419579e58d7183136184b0071faac3a3","description":"\n\t\n\t\t\n\t\n\t\n\t\tUpload your Adapter Oracle and Parity results\n\t\n\nThis dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.\nadapters/\n  └── {adapter_name}/\n      ├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE. \n      ├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.\n      ├──… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/parity-experiments.","downloads":9636,"tags":["region:us"],"createdAt":"2025-12-24T23:23:33.000Z","key":""},{"_id":"694e1faf785ac06cb1c0e866","id":"harborframework/harbor-datasets","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-04-23T19:17:39.000Z","likes":0,"trendingScore":0,"private":false,"sha":"bae16d3bc8b10aab97018e8327cb4db4806d6033","downloads":8945,"tags":["license:apache-2.0","region:us"],"createdAt":"2025-12-26T05:39:59.000Z","key":""},{"_id":"699e5201a11c4e3856c618a7","id":"harborframework/terminal-bench-science-lfs","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-08-22T21:58:18.000Z","likes":0,"trendingScore":0,"private":false,"sha":"d48ce2a26e4dcd525c9e8ceb8dd2b60652e178cd","description":"\n\t\n\t\t\n\t\n\t\n\t\tTerminal-Bench-Science — task input mirror\n\t\n\nLarge input files for Terminal-Bench-Science\ntasks, which cannot be committed to git. Tasks pull from here at container build\ntime, pinned to a commit SHA and verified against a checksum file that ships in\nthe task directory.\nOne top-level prefix per task; everything lives under <task-name>/input/.\n\n\t\n\t\t\n\t\n\t\n\t\tBenchmark contamination canary\n\t\n\nThis dataset is benchmark material. If you are assembling a training corpus,\nexclude it.… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-science-lfs.","downloads":9115,"tags":["license:apache-2.0","region:us","benchmark","terminal-bench-science"],"createdAt":"2026-02-25T01:36:01.000Z","key":""},{"_id":"69f94db8ab7325c216725209","id":"harborframework/harbor-mix","author":"harborframework","disabled":false,"gated":"auto","lastModified":"2026-05-07T10:48:10.000Z","likes":3,"trendingScore":0,"private":false,"sha":"c9696fd51d5b2ddf4a84ee24c41d40816f0f43af","description":"\n\t\n\t\t\n\t\tHarbor-Mix\n\t\n\nHarbor-Mix is a curated meta-dataset of 100 difficult, diverse, and high-quality agentic evaluation tasks selected from the Harbor Adapters benchmark pool. It is designed to preserve broad signal from large-scale agent evaluations while being substantially cheaper to run than a full multi-benchmark sweep.\n\n\t\n\t\t\n\t\tWhat Is Included\n\t\n\nThis repository contains the 100 task directories, flattened at the repository root.\nEach task directory contains:\n\ninstruction.md: the… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/harbor-mix.","downloads":213,"tags":["task_categories:text-generation","task_categories:reinforcement-learning","license:cc-by-4.0","size_categories:n<1K","region:us","agent-benchmark","llm-agents","agentic-evaluation","harbor","tool-use","software-engineering","scientific-reasoning","llm-as-a-judge"],"createdAt":"2026-05-05T01:54:00.000Z","key":""},{"_id":"6a8d264f2a2cb9f2717ef994","id":"harborframework/terminal-bench-lfs","author":"harborframework","disabled":false,"gated":false,"lastModified":"2026-08-25T05:21:37.000Z","likes":0,"trendingScore":0,"private":false,"sha":"f01fff049773b5a3141538a1f3cb3dd1beadae82","downloads":1943,"tags":["region:us"],"createdAt":"2026-08-25T05:21:19.000Z","key":""}]