Harbor Hub
Find and share Harbor tasks for training and evaluation.
| Dataset | Tasks |
|---|---|
frontier-bench/frontier-bench The frontier is wide and dynamic. Frontier-Bench measures general agent capabilities on a diverse set of difficult tasks. | 74 |
terminal-bench/terminal-bench-2-1 Version 2.1 of Terminal-Bench, a benchmark for evaluating agents in terminal environments. | 89 |
harbor-index/harbor-index-1.0 Harbor Index is an 82-task benchmark for agentic evaluation. It was distilled from more than 6,000 candidate tasks through repeated model runs, automated broken-task identification, human audits, and reward hacking supervision. | 82 |
android-bench/android-bench A dataset of 100 Android software engineering tasks derived from real-world repositories. | 100 |
scale-ai/swe-bench-pro | 731 |
swe-bench/swe-bench-verified | 500 |
Publish your first dataset
Add the Harbor skill to your coding agent, then run /publish.
npx skills add harbor-framework/harbor --skill publish/publish