Loading datasets…
Loading datasets…
Loading datasets…
Organization collection
Datasets published by this organization.
Organization inventory
Private datasets are available only to members of this organization.
| Dataset | AccessVisibility | Tasks |
|---|---|---|
silrlab/long-horizon-terminal-bench Long-Horizon Terminal-Bench: 46-task benchmark measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Hidden rebuild-from-artifact verifiers; self-reported progress does not count. | Public | 46 |
silrlab/long-horizon-terminal-bench-2.0 Long-Horizon Terminal-Bench 2.0: 77-task benchmark measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. The 46 original tasks plus 31 new ones, 30 of them reconstruction tasks. Hidden rebuild-from-artifact verifiers; self-reported progress does not count. | Public | 77 |
2 datasets
Organization collection
Datasets published by this organization.
Organization inventory
Private datasets are available only to members of this organization.
| Dataset | AccessVisibility | Tasks |
|---|---|---|
silrlab/long-horizon-terminal-bench Long-Horizon Terminal-Bench: 46-task benchmark measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Hidden rebuild-from-artifact verifiers; self-reported progress does not count. | Public | 46 |
silrlab/long-horizon-terminal-bench-2.0 Long-Horizon Terminal-Bench 2.0: 77-task benchmark measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. The 46 original tasks plus 31 new ones, 30 of them reconstruction tasks. Hidden rebuild-from-artifact verifiers; self-reported progress does not count. | Public | 77 |
2 datasets