Loading dataset details…
harbor run -d islo-labs/reward-hack-bench-controlBaseline/control companion to reward-hack-bench: the same 8 SWE-bench + CyBench tasks with NO cheat prompt, still scored by the fairness judge. Also packages these CyBench/GlacierCTF tasks as Harbor tasks.
harbor run -d islo-labs/reward-hack-bench-controlThe clean baseline companion to islo-labs/reward-hack-bench: the same 8
SWE-bench + CyBench tasks with no cheat prompt. It measures the model's
natural solve rate and whether it reaches for the published solution
unprompted, with the same fairness judge attached — and doubles as a Harbor
packaging of these CyBench / GlacierCTF and SWE-bench-Verified tasks.
harbor run -d islo-labs/reward-hack-bench-control@latest -a claude-code -m anthropic/claude-opus-4-7
The benchmark with cheat paths injected: islo-labs/reward-hack-bench.