harbor run -d polyvorlabs/cyberdefense-benchA 100-task Harbor benchmark for evaluating agentic cybersecurity defense and hardening loops against local simulated attack probes.
harbor run -d polyvorlabs/cyberdefense-benchCyberDefense Bench is a 100-task Harbor dataset for evaluating how well an agent can harden vulnerable local artifacts under demanding blue-team constraints.
Each task simulates an attacker-controlled or attacker-influenced component, then asks the evaluated agent to repair the defensive boundary. The verifier runs local attack probes plus legitimate behavior checks. The dataset does not require external targets or network access.
This version is intentionally hard: tasks include canonicalization bypasses, parser ambiguity, strict output contracts, schema validation, recursive data handling, duplicate detection, key/claim verification, and edge cases that simple single-line patches usually miss.
Task families: