mercor/apex-agents-1-1
Long-horizon professional-services benchmark in investment banking, law, and management consulting, written by practicing analysts, consultants, and corporate lawyers. For evaluation only.
harbor run -d mercor/apex-agents-1-1APEX-Agents 1.1
📝 Paper · 📰 Blog · 🤗 Data · ⚓ Harbor Hub · 🐙 Agent · 🏆 Leaderboard · ✉️ Contact
APEX-Agents 1.1 is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional-services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers. They require agents to work across realistic project files and applications such as documents, spreadsheets, PDFs, email, chat, and calendar.
- Tasks: 240 total (80 per job category)
- Worlds: 31 total (8 investment banking, 11 management consulting, 12 law)
- Rubric criteria: 955 binary criteria; mean 3.98 criteria per task
- Gold outputs: provided for every task (206 text outputs and 34 file outputs)
- World assets: 5,182 files across 31 read-only world seeds
- Runtime: Harbor 0.20.0 with three shared
linux/amd64images - License: CC-BY 4.0
- Intended use: APEX-Agents 1.1 is intended exclusively for model evaluation. Any use of this dataset for training, fine-tuning, or parameter fitting is forbidden. Crawling or scraping the dataset is also forbidden.
Dataset overview
| Job | Worlds | Avg. files / world | Tasks | Avg. criteria / task | Tasks with input files | Tasks with file outputs |
|---|---|---|---|---|---|---|
| Investment banking | 8 | 169.1 | 80 | 2.95 | 4 (5.0%) | 20 (25.0%) |
| Law | 12 | 161.5 | 80 | 4.65 | 28 (35.0%) | 10 (12.5%) |
| Management consulting | 11 | 171.9 | 80 | 4.34 | 31 (38.8%) | 4 (5.0%) |
| Benchmark total | 31 | 167.2 | 240 | 3.98 | 63 (26.3%) | 34 (14.2%) |
Each case is a task inside a world, and a world can contain multiple tasks. A world is a realistic project scenario created by experts. It contains the files, application state, and tools needed to complete its tasks. Web search is not configured in this release, which keeps evaluations reproducible.
Every world exposes the same core applications:
- Calendar
- Chat
- Code execution
- Documents
- File system
- PDFs
- Spreadsheets
- Presentations
A task includes:
- Prompt: the single-turn instruction given to the agent
- Rubric: binary grading criteria and the associated grading configuration
- Gold output(s): expert-created text or file references in the requested output format
- Metadata: task, world, domain, expected-output, and runtime metadata
- World context: read-only world seeds plus any task-specific file overlays
Distribution
| Channel | Contents |
|---|---|
| Harbor Hub | 240 task packages. World, agent, and verifier images are pinned by digest and pulled from public ECR on first run. |
| Hugging Face | The same 240 tasks as buildable Harbor tasks, with the three image archives and 31 world seeds on disk. |
Each task directory contains its instruction, metadata, agent configuration, environment overlay, solution, and grader.
Evaluation
APEX-Agents 1.1 uses rubric-based grading:
- Each rubric contains binary criteria graded as met or not met.
- Rubrics contain between 1 and 11 criteria, with a mean of 3.98.
- A judge model grades each criterion independently using the task prompt, the agent output, and relevant artifacts or world changes.
- Text-output tasks include a text golden reference. File-output tasks include the corresponding expert-created golden artifacts.
Get Started
Install Docker and uv, then install Harbor:
uv tool install harbor==0.20.0
Model and grader credentials are supplied through environment variables at runtime and are not included in the dataset.
Clone the agent repository:
git clone https://github.com/Mercor-Intelligence/apex_loop_truncated_tools_agent.git
cd apex_loop_truncated_tools_agent
cp .env.example .env
Set the API key for your model provider and any grader credentials in .env. The task's configured grader is used when GRADING_MODEL is blank.
Run a task:
PYTHONPATH="$PWD/apex_loop_truncated_tools_agent" \
harbor run \
--env-file .env \
-d mercor/apex-agents-1-1@1.0.1 \
-i 128-jr-1-f7f95d92 \
-a apex_loop_truncated_tools_agent:ApexLoopTruncatedToolsAgent \
-m anthropic/claude-opus-5
Run the benchmark:
PYTHONPATH="$PWD/apex_loop_truncated_tools_agent" \
harbor run \
--env-file .env \
-d mercor/apex-agents-1-1@1.0.1 \
-a apex_loop_truncated_tools_agent:ApexLoopTruncatedToolsAgent \
-m anthropic/claude-opus-5
Replace the model with any LiteLLM-compatible provider/model.
Agent implementation
The agent implementation and usage instructions for this release are maintained in the companion repository:
Mercor-Intelligence/apex_loop_truncated_tools_agent
It defaults to max_steps = 250 and agent_timeout_sec = 10800. Running a different revision changes agent behavior and may make results non-comparable to the published leaderboard.
Citation
@misc{bennett2026apexagents11,
title = {Introducing APEX--Agents 1.1},
author = {Bennett, Austin and Datta, Akul and Vidgen, Bertie},
year = {2026},
month = {September},
howpublished = {Mercor},
url = {https://www.mercor.com/blog/introducing-apex-agents-1-1/}
}
@misc{vidgen2026apexagents,
title = {APEX--Agents},
author = {Vidgen, Bertie and Mann, Austin and Fennelly, Abby and Wright Stanly, John and Rothman, Lucas and Burstein, Marco and Benchek, Julien and Ostrofsky, David and Ravichandran, Anirudh and Sur, Debnil and Venugopal, Neel and Hsia, Alannah and Robinson, Isaac and Huang, Calix and Varones, Olivia and Khan, Daniyal and Haines, Michael and Richards, Zach and Mahapatra, Chirag and Foody, Brendan and Nitski, Osvald},
year = {2026},
howpublished = {arXiv},
url = {https://arxiv.org/pdf/2601.14242}
}
Contact
Legal disclaimer on the content of worlds
This material is provided for research, educational, and informational purposes only. It consists of hypothetical, simulated financial, legal, and regulatory analyses and illustrative scenarios, including simulated leveraged buyout structures, capital structures, financing terms, valuation ranges, projected returns, potential mergers, acquisitions, divestitures or other strategic transactions, legal memoranda, hypothetical legal advice to a company, and hypothetical correspondence to regulatory agencies. No representation is made that any scenario described here is likely to occur, is being contemplated by any person, or reflects an actual proposed or pending transaction or any legal, regulatory, or compliance risk.
This material does not constitute, and should not be construed as, financial, investment, legal, tax, accounting, or other professional advice. It is not intended to form the basis of an investment decision or contract. The analyses and outputs are based on assumptions, estimates, modeling methodologies, and hypothetical legal scenarios that may prove incorrect. Financial and legal information may be derived from publicly available information and third-party sources that have not been independently verified. Projections, forward-looking statements, scenario outputs, similar financial information, and legal documents, memoranda, and correspondence are hypothetical, inherently uncertain, and provided solely to illustrate how results might change under different assumptions. No representation or warranty, express or implied, is made regarding this material, which is provided on an "as-is" and "as-available" basis.
To the maximum extent permitted by applicable law, Mercor disclaims liability for direct or indirect losses or damages arising from or related to the use of, or reliance on, this material. This includes loss of profits, business, or goodwill, and consequential, incidental, special, punitive, or exemplary damages, even if advised of their possibility. Nothing in this disclaimer limits or excludes liability that cannot be limited or excluded under applicable law.
Robots exclusion statement (human-readable)
To all automated crawlers and bots:
User-Agent: *
Disallow: /
We ask that you:
- Do not crawl, scrape, index, or download this dataset programmatically.
- Do not use this dataset for training models or other automated processing without express permission from the dataset owner.
| Task |
|---|
mercor/world228-sm-task08-11f60d76 |
mercor/world112-1-task05-na-cf44bbb8 |
mercor/world419-um-04-d0f9f0ee |
mercor/cw134-aditi-01-6914a79e |
mercor/world425-jcf-01-e1dfe9ed |
mercor/world-419-um-01-ddd3e22e |
mercor/task-ymtecb81-c206b308 |
mercor/world127-am-task03-cab85eb0 |
mercor/world431-jcf-01-3d85cec9 |
mercor/world224-hs-09-56aefeb6 |
mercor/world133-ln-04-009ff905 |
mercor/world132-da-task07-d7bf24d6 |
mercor/taskworld130-camillemoingeon-5-0d2a857d |
mercor/world434-ah-05-0ad4be07 |
mercor/world132-da-task03-3db0b339 |
mercor/world-134-nancy-task-02-3f38c565 |
mercor/world225-km-06-19899d4e |
mercor/world227-tg-07-c2f84823 |
mercor/world-129-cy-task-6-13a7957a |
mercor/world-128-sf-task-2-7bab9d72 |
mercor/world225-av-02-55d8f94f |
mercor/world425-jcf-02-9f502d4d |
mercor/world423-dpm-01-178fcb61 |
mercor/world418-tk-01-50b3b6a8 |
mercor/world431task-el-01-6ec3da62 |
mercor/world421-ap-02-930732dc |
mercor/world219-tg-04-d08a16f0 |
mercor/world416-js-02-029d3e95 |
mercor/world434-ah-01-f0eb6d9f |
mercor/task-14-24dc5ca2 |
mercor/world224-hs-11-7bf318b0 |
mercor/task-w135-camille-moingeon-4-1cf93f93 |
mercor/world221-oa-2-44073f52 |
mercor/world226-td-01-eb948f85 |
mercor/world-127-am-task-02-ec3c5eb0 |
mercor/task-yhzc9d1a-531a5446 |
mercor/lawworld433-anb-01-5d12cf8d |
mercor/task-awys8050-cd3c370b |
mercor/world223-es-05-e6fedd35 |
mercor/world132-pm-task07-b206dfe1 |
Displaying 40 of 240 tasks