harbor run -d terminal-bench/terminal-bench-cpu-onlyharbor run -d terminal-bench/terminal-bench-cpu-onlyThe subset of Terminal-Bench tasks.
harbor run -d terminal-bench/terminal-bench-cpu-onlyThe subset of Terminal-Bench tasks.
harbor run -d terminal-bench/terminal-bench-cpu-onlyThis is a subset of Terminal-Bench that does not include the (3) tasks which require GPUs.
Terminal-Bench measures general agent capabilities on a diverse set of difficult tasks and is continuously updated with new and improved tasks.
First, install the Harbor CLI.
Tasks have been validated on both Modal and Daytona.
uv tool install "harbor[modal,daytona]"
# or pip install "harbor[modal,daytona]"
And then run:
harbor run -d terminal-bench/terminal-bench-cpu-only \
--agent claude-code \
--model anthropic/claude-fable-5 \
--n-concurrent 100 \
--env modal
To upload results to Harbor Hub, include the --upload flag or run harbor upload "<job/path>".
See CONTRIBUTING.md for how to propose, implement, and submit new tasks.
We also encourage the community to reach out when they find bugs in tasks.
This is a subset of Terminal-Bench that does not include the (3) tasks which require GPUs.
Terminal-Bench measures general agent capabilities on a diverse set of difficult tasks and is continuously updated with new and improved tasks.
First, install the Harbor CLI.
Tasks have been validated on both Modal and Daytona.
uv tool install "harbor[modal,daytona]"
# or pip install "harbor[modal,daytona]"
And then run:
harbor run -d terminal-bench/terminal-bench-cpu-only \
--agent claude-code \
--model anthropic/claude-fable-5 \
--n-concurrent 100 \
--env modal
To upload results to Harbor Hub, include the --upload flag or run harbor upload "<job/path>".
See CONTRIBUTING.md for how to propose, implement, and submit new tasks.
We also encourage the community to reach out when they find bugs in tasks.