Harbor Hub

Find and share Harbor tasks for training and evaluation.

DatasetTasks
frontier-bench/frontier-bench

The frontier is wide and dynamic. Frontier-Bench measures general agent capabilities on a diverse set of difficult tasks.

74
terminal-bench/terminal-bench-2-1

Version 2.1 of Terminal-Bench, a benchmark for evaluating agents in terminal environments.

89
harbor-index/harbor-index-1.0

Harbor Index is an 82-task benchmark for agentic evaluation. It was distilled from more than 6,000 candidate tasks through repeated model runs, automated broken-task identification, human audits, and reward hacking supervision.

82
android-bench/android-bench

A dataset of 100 Android software engineering tasks derived from real-world repositories.

100
scale-ai/swe-bench-pro
731
swe-bench/swe-bench-verified
500

Publish your first dataset

Add the Harbor skill to your coding agent, then run /publish.

1. Install the Harbor skill
npx skills add harbor-framework/harbor --skill publish
2. Run /publish in your agent
/publish