scale-ai/swe-bench-pro

Published 4/15/2026 by

harbor run -d scale-ai/swe-bench-pro

SWE-Bench Pro is a benchmark designed to provide a rigorous and realistic evaluation of AI agents for software engineering. It was developed to address several limitations in existing benchmarks by tackling four key challenges:

  1. Data Contamination: Models have likely seen the evaluation code during training, making it hard to know if they are problem-solving or recalling a memorized solution.
  2. Limited Task Diversity: Many benchmarks fail to capture the full spectrum of real-world software challenges and instead focus on simple utility libraries.
  3. Oversimplified Problems: Ambiguous or underspecified issues are often removed from benchmarks, which doesn't reflect a real developer's workflow.
  4. Unreliable and Irreproducible Testing: Inconsistent setups make it difficult to know if a solution truly works or if the environment is just configured incorrectly.

SWE-Bench Pro addresses these gaps by sourcing tasks from diverse and complex codebases, including consumer applications, B2B services, and developer tools. To reduce contamination risk, the public and held-out OSS subsets use strong copyleft licenses (e.g., GPL). The private subset consists of proprietary codebases from startup partners.

The benchmark is significantly more challenging than its predecessors; top models score around 23% on the SWE-Bench Pro public set, compared to 70%+ on SWE-Bench Verified. This provides a more accurate measure of an agent’s true problem-solving capabilities in environments that mirror professional software development.

Read the paper here: https://scale.com/research/swe_bench_pro

Check out the GitHub and view trajectories here: https://docent.transluce.org/dashboard/032fb63d-4992-4bfc-911d-3b7dafcb931f

(https://labs.scale.com/leaderboard/swe_bench_pro_public)

Task
scale-ai/instance_protonmail__webclients-d3e513044d299d04e509bf8c0f4e73d812030246
scale-ai/instance_qutebrowser__qutebrowser-36ade4bba504eb96f05d32ceab9972df7eb17bcc-v2ef375ac784985212b1805e1d0431dc8f1b3c171
scale-ai/instance_future-architect__vuls-3c1489e588dacea455ccf4c352a3b1006902e2d4
scale-ai/instance_protonmail__webclients-5e815cfa518b223a088fa9bb232a5fc90ab15691
scale-ai/instance_ansible__ansible-9142be2f6cabbe6597c9254c5bb9186d17036d55-v0f01c69f1e2528b935359cfe578530722bca2c59
scale-ai/instance_flipt-io__flipt-b3cd920bbb25e01fdb2dab66a5a913363bc62f6c
scale-ai/instance_element-hq__element-web-f0359a5c180b8fec4329c77adcf967c8d3b7b787-vnan
scale-ai/instance_protonmail__webclients-428cd033fede5fd6ae9dbc7ab634e010b10e4209
scale-ai/instance_navidrome__navidrome-55730514ea59d5f1d0b8e3f8745569c29bdbf7b4
scale-ai/instance_flipt-io__flipt-492cc0b158200089dceede3b1aba0ed28df3fb1d
scale-ai/instance_protonmail__webclients-08bb09914d0d37b0cd6376d4cab5b77728a43e7b
scale-ai/instance_qutebrowser__qutebrowser-e34dfc68647d087ca3175d9ad3f023c30d8c9746-v363c8a7e5ccdf6968fc7ab84a2053ac78036691d
scale-ai/instance_ansible__ansible-5e369604e1930b1a2e071fecd7ec5276ebd12cb1-v0f01c69f1e2528b935359cfe578530722bca2c59
scale-ai/instance_qutebrowser__qutebrowser-7b603dd6bf195e3e723ce08ff64a82b406e3f6b6-v9f8e9d96c85c85a605e382f1510bd08563afc566
scale-ai/instance_flipt-io__flipt-c8d71ad7ea98d97546f01cce4ccb451dbcf37d3b
scale-ai/instance_nodebb__nodebb-8168c6c40707478f71b8af60300830fe554c778c-vf2cf3cbd463b7ad942381f1c6d077626485a1e9e
scale-ai/instance_flipt-io__flipt-7161f7b876773a911afdd804b281e52681cb7321
scale-ai/instance_future-architect__vuls-be7b9114cc9545e68fb0ee7bc63d7ec53d1a00ad
scale-ai/instance_qutebrowser__qutebrowser-3fd8e12949b8feda401930574facf09dd4180bba
scale-ai/instance_element-hq__element-web-582a1b093fc0b77538052f45cbb9c7295f991b51-vnan
scale-ai/instance_gravitational__teleport-007235446f85b1cbaef92664c3b3867517250f21
scale-ai/instance_future-architect__vuls-aaea15e516ece43978cf98e09e52080478b1d39f
scale-ai/instance_protonmail__webclients-5d2576632037d655c3b6a28e98cd157f7e9a5ce1
scale-ai/instance_flipt-io__flipt-ebb3f84c74d61eee4d8c6875140b990eee62e146
scale-ai/instance_qutebrowser__qutebrowser-fec187c2cb53d769c2682b35ca77858a811414a8-v363c8a7e5ccdf6968fc7ab84a2053ac78036691d
scale-ai/instance_navidrome__navidrome-8e640bb8580affb7e0ea6225c0bbe240186b6b08
scale-ai/instance_qutebrowser__qutebrowser-bedc9f7fadf93f83d8dee95feeecb9922b6f063f-v2ef375ac784985212b1805e1d0431dc8f1b3c171
scale-ai/instance_protonmail__webclients-09fcf0dbdb87fa4f4a27700800ee4a3caed8b413
scale-ai/instance_navidrome__navidrome-874b17b8f614056df0ef021b5d4f977341084185
scale-ai/instance_gravitational__teleport-2b15263e49da5625922581569834eec4838a9257-vee9b09fb20c43af7e520f57e9239bbcf46b7113d
scale-ai/instance_ansible__ansible-34db57a47f875d11c4068567b9ec7ace174ec4cf-v1055803c3a812189a1133297f7f5468579283f86

Displaying 31 of 731 tasks