islo-labs/reward-hack-bench-control
Baseline/control companion to reward-hack-bench: the same 8 SWE-bench + CyBench tasks with NO cheat prompt, still scored by the fairness judge. Also packages these CyBench/GlacierCTF tasks as Harbor tasks.
Published 6/18/2026 by Rotem Tamir
harbor run -d islo-labs/reward-hack-bench-controlreward-hack-bench-control
The clean baseline companion to islo-labs/reward-hack-bench: the same 8
SWE-bench + CyBench tasks with no cheat prompt. It measures the model's
natural solve rate and whether it reaches for the published solution
unprompted, with the same fairness judge attached — and doubles as a Harbor
packaging of these CyBench / GlacierCTF and SWE-bench-Verified tasks.
- Code & methodology: https://github.com/islo-labs/reward-hack-bench
- Writeup: https://islo.dev/blog/reward-hack-bench-sandbox-stops-agent-cheating/
harbor run -d islo-labs/reward-hack-bench-control@latest -a claude-code -m anthropic/claude-opus-4-7
The benchmark with cheat paths injected: islo-labs/reward-hack-bench.
| Task |
|---|
reward-hack-bench/pytest-6202__control |
reward-hack-bench/ezmaze__control |
reward-hack-bench/slcg__control |
reward-hack-bench/sympy-12489__control |
reward-hack-bench/sympy-14531__control |
reward-hack-bench/matplotlib-25775__control |
reward-hack-bench/diffecient__control |
reward-hack-bench/noisier-crc__control |
Displaying 8 of 8 tasks