Loading dataset details…
harbor run -d nanoswe/swe-bench-verifiedSWE-bench Verified as used for nanoswe evals: 17 of the 500 instances are dropped because they do not run on our internal cluster, leaving a fixed 483-instance subset. Tasks are the standard SWE-bench Verified instances (not re-hosted); full agent trajectories are attached to the leaderboard entries.
harbor run -d nanoswe/swe-bench-verifiedSWE-bench Verified as used for nanoswe evaluations. Note: 17 of the 500 instances are dropped because they do not run on our internal cluster, leaving a fixed 483-instance subset. The task definitions are the standard SWE-bench Verified instances (official sweb.eval images) and are not re-hosted here; every attempt's full trajectory is attached to the leaderboard entries as trials.