harbor run -d mercor/apex-agents-1-1Long-horizon professional-services benchmark in investment banking, law, and management consulting, written by practicing analysts, consultants, and corporate lawyers. For evaluation only.
harbor run -d mercor/apex-agents-1-1📝 Paper · 📰 Blog · 🤗 Data · ⚓ Harbor Hub · 🐙 Agent · 🏆 Leaderboard · ✉️ Contact
APEX-Agents 1.1 is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional-services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers. They require agents to work across realistic project files and applications such as documents, spreadsheets, PDFs, email, chat, and calendar.
linux/amd64 images| Job | Worlds | Avg. files / world | Tasks | Avg. criteria / task | Tasks with input files | Tasks with file outputs |
|---|---|---|---|---|---|---|
| Investment banking | 8 | 169.1 | 80 | 2.95 | 4 (5.0%) | 20 (25.0%) |
| Law | 12 | 161.5 | 80 | 4.65 | 28 (35.0%) | 10 (12.5%) |
| Management consulting | 11 | 171.9 | 80 | 4.34 | 31 (38.8%) | 4 (5.0%) |
| Benchmark total | 31 | 167.2 | 240 | 3.98 | 63 (26.3%) | 34 (14.2%) |
Each case is a task inside a world, and a world can contain multiple tasks. A world is a realistic project scenario created by experts. It contains the files, application state, and tools needed to complete its tasks. Web search is not configured in this release, which keeps evaluations reproducible.
Every world exposes the same core applications:
A task includes:
| Channel | Contents |
|---|---|
| Harbor Hub | 240 task packages. World, agent, and verifier images are pinned by digest and pulled from public ECR on first run. |
| Hugging Face | The same 240 tasks as buildable Harbor tasks, with the three image archives and 31 world seeds on disk. |
Each task directory contains its instruction, metadata, agent configuration, environment overlay, solution, and grader.
APEX-Agents 1.1 uses rubric-based grading:
Install Docker and uv, then install Harbor:
uv tool install harbor==0.20.0
Model and grader credentials are supplied through environment variables at runtime and are not included in the dataset.
Clone the agent repository:
git clone https://github.com/Mercor-Intelligence/apex_loop_truncated_tools_agent.git
cd apex_loop_truncated_tools_agent
cp .env.example .env
Set the API key for your model provider and any grader credentials in .env. The task's configured grader is used when GRADING_MODEL is blank.
Run a task:
PYTHONPATH="$PWD/apex_loop_truncated_tools_agent" \
harbor run \
--env-file .env \
-d mercor/apex-agents-1-1@1.0.1 \
-i 128-jr-1-f7f95d92 \
-a apex_loop_truncated_tools_agent:ApexLoopTruncatedToolsAgent \
-m anthropic/claude-opus-5
Run the benchmark:
PYTHONPATH="$PWD/apex_loop_truncated_tools_agent" \
harbor run \
--env-file .env \
-d mercor/apex-agents-1-1@1.0.1 \
-a apex_loop_truncated_tools_agent:ApexLoopTruncatedToolsAgent \
-m anthropic/claude-opus-5
Replace the model with any LiteLLM-compatible provider/model.
The agent implementation and usage instructions for this release are maintained in the companion repository:
Mercor-Intelligence/apex_loop_truncated_tools_agent
The loop harness runs 100 turns (max_steps = 100) with agent_timeout_sec = 10800, matching the leaderboard runs. Running a different revision or step budget changes agent behavior and may make results non-comparable to the published leaderboard.
@misc{bennett2026apexagents11,
title = {Introducing APEX--Agents 1.1},
author = {Bennett, Austin and Datta, Akul and Vidgen, Bertie},
year = {2026},
month = {September},
howpublished = {Mercor},
url = {https://www.mercor.com/blog/introducing-apex-agents-1-1/}
}
@misc{vidgen2026apexagents,
title = {APEX--Agents},
author = {Vidgen, Bertie and Mann, Austin and Fennelly, Abby and Wright Stanly, John and Rothman, Lucas and Burstein, Marco and Benchek, Julien and Ostrofsky, David and Ravichandran, Anirudh and Sur, Debnil and Venugopal, Neel and Hsia, Alannah and Robinson, Isaac and Huang, Calix and Varones, Olivia and Khan, Daniyal and Haines, Michael and Richards, Zach and Mahapatra, Chirag and Foody, Brendan and Nitski, Osvald},
year = {2026},
howpublished = {arXiv},
url = {https://arxiv.org/pdf/2601.14242}
}
This material is provided for research, educational, and informational purposes only. It consists of hypothetical, simulated financial, legal, and regulatory analyses and illustrative scenarios, including simulated leveraged buyout structures, capital structures, financing terms, valuation ranges, projected returns, potential mergers, acquisitions, divestitures or other strategic transactions, legal memoranda, hypothetical legal advice to a company, and hypothetical correspondence to regulatory agencies. No representation is made that any scenario described here is likely to occur, is being contemplated by any person, or reflects an actual proposed or pending transaction or any legal, regulatory, or compliance risk.
This material does not constitute, and should not be construed as, financial, investment, legal, tax, accounting, or other professional advice. It is not intended to form the basis of an investment decision or contract. The analyses and outputs are based on assumptions, estimates, modeling methodologies, and hypothetical legal scenarios that may prove incorrect. Financial and legal information may be derived from publicly available information and third-party sources that have not been independently verified. Projections, forward-looking statements, scenario outputs, similar financial information, and legal documents, memoranda, and correspondence are hypothetical, inherently uncertain, and provided solely to illustrate how results might change under different assumptions. No representation or warranty, express or implied, is made regarding this material, which is provided on an "as-is" and "as-available" basis.
To the maximum extent permitted by applicable law, Mercor disclaims liability for direct or indirect losses or damages arising from or related to the use of, or reliance on, this material. This includes loss of profits, business, or goodwill, and consequential, incidental, special, punitive, or exemplary damages, even if advised of their possibility. Nothing in this disclaimer limits or excludes liability that cannot be limited or excluded under applicable law.
To all automated crawlers and bots:
User-Agent: *
Disallow: /
We ask that you: