mercor/apex-agents-1-1

Long-horizon professional-services benchmark in investment banking, law, and management consulting, written by practicing analysts, consultants, and corporate lawyers. For evaluation only.

harbor run -d mercor/apex-agents-1-1

APEX-Agents 1.1

📝 Paper · 📰 Blog · 🤗 Data · ⚓ Harbor Hub · 🐙 Agent · 🏆 Leaderboard · ✉️ Contact

APEX-Agents 1.1 is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional-services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers. They require agents to work across realistic project files and applications such as documents, spreadsheets, PDFs, email, chat, and calendar.

  • Tasks: 240 total (80 per job category)
  • Worlds: 31 total (8 investment banking, 11 management consulting, 12 law)
  • Rubric criteria: 955 binary criteria; mean 3.98 criteria per task
  • Gold outputs: provided for every task (206 text outputs and 34 file outputs)
  • World assets: 5,182 files across 31 read-only world seeds
  • Runtime: Harbor 0.20.0 with three shared linux/amd64 images
  • License: CC-BY 4.0
  • Intended use: APEX-Agents 1.1 is intended exclusively for model evaluation. Any use of this dataset for training, fine-tuning, or parameter fitting is forbidden. Crawling or scraping the dataset is also forbidden.

Dataset overview

Job Worlds Avg. files / world Tasks Avg. criteria / task Tasks with input files Tasks with file outputs
Investment banking 8 169.1 80 2.95 4 (5.0%) 20 (25.0%)
Law 12 161.5 80 4.65 28 (35.0%) 10 (12.5%)
Management consulting 11 171.9 80 4.34 31 (38.8%) 4 (5.0%)
Benchmark total 31 167.2 240 3.98 63 (26.3%) 34 (14.2%)

Each case is a task inside a world, and a world can contain multiple tasks. A world is a realistic project scenario created by experts. It contains the files, application state, and tools needed to complete its tasks. Web search is not configured in this release, which keeps evaluations reproducible.

Every world exposes the same core applications:

  • Calendar
  • Chat
  • Code execution
  • Documents
  • File system
  • Mail
  • PDFs
  • Spreadsheets
  • Presentations

A task includes:

  • Prompt: the single-turn instruction given to the agent
  • Rubric: binary grading criteria and the associated grading configuration
  • Gold output(s): expert-created text or file references in the requested output format
  • Metadata: task, world, domain, expected-output, and runtime metadata
  • World context: read-only world seeds plus any task-specific file overlays

Distribution

Channel Contents
Harbor Hub 240 task packages. World, agent, and verifier images are pinned by digest and pulled from public ECR on first run.
Hugging Face The same 240 tasks as buildable Harbor tasks, with the three image archives and 31 world seeds on disk.

Each task directory contains its instruction, metadata, agent configuration, environment overlay, solution, and grader.

Evaluation

APEX-Agents 1.1 uses rubric-based grading:

  • Each rubric contains binary criteria graded as met or not met.
  • Rubrics contain between 1 and 11 criteria, with a mean of 3.98.
  • A judge model grades each criterion independently using the task prompt, the agent output, and relevant artifacts or world changes.
  • Text-output tasks include a text golden reference. File-output tasks include the corresponding expert-created golden artifacts.

Get Started

Install Docker and uv, then install Harbor:

uv tool install harbor==0.20.0

Model and grader credentials are supplied through environment variables at runtime and are not included in the dataset.

Clone the agent repository:

git clone https://github.com/Mercor-Intelligence/apex_loop_truncated_tools_agent.git
cd apex_loop_truncated_tools_agent
cp .env.example .env

Set the API key for your model provider and any grader credentials in .env. The task's configured grader is used when GRADING_MODEL is blank.

Run a task:

PYTHONPATH="$PWD/apex_loop_truncated_tools_agent" \
harbor run \
  --env-file .env \
  -d mercor/apex-agents-1-1@1.0.1 \
  -i 128-jr-1-f7f95d92 \
  -a apex_loop_truncated_tools_agent:ApexLoopTruncatedToolsAgent \
  -m anthropic/claude-opus-5

Run the benchmark:

PYTHONPATH="$PWD/apex_loop_truncated_tools_agent" \
harbor run \
  --env-file .env \
  -d mercor/apex-agents-1-1@1.0.1 \
  -a apex_loop_truncated_tools_agent:ApexLoopTruncatedToolsAgent \
  -m anthropic/claude-opus-5

Replace the model with any LiteLLM-compatible provider/model.

Agent implementation

The agent implementation and usage instructions for this release are maintained in the companion repository:

Mercor-Intelligence/apex_loop_truncated_tools_agent

It defaults to max_steps = 250 and agent_timeout_sec = 10800. Running a different revision changes agent behavior and may make results non-comparable to the published leaderboard.

Citation

@misc{bennett2026apexagents11,
  title        = {Introducing APEX--Agents 1.1},
  author       = {Bennett, Austin and Datta, Akul and Vidgen, Bertie},
  year         = {2026},
  month        = {September},
  howpublished = {Mercor},
  url          = {https://www.mercor.com/blog/introducing-apex-agents-1-1/}
}

@misc{vidgen2026apexagents,
  title        = {APEX--Agents},
  author       = {Vidgen, Bertie and Mann, Austin and Fennelly, Abby and Wright Stanly, John and Rothman, Lucas and Burstein, Marco and Benchek, Julien and Ostrofsky, David and Ravichandran, Anirudh and Sur, Debnil and Venugopal, Neel and Hsia, Alannah and Robinson, Isaac and Huang, Calix and Varones, Olivia and Khan, Daniyal and Haines, Michael and Richards, Zach and Mahapatra, Chirag and Foody, Brendan and Nitski, Osvald},
  year         = {2026},
  howpublished = {arXiv},
  url          = {https://arxiv.org/pdf/2601.14242}
}

Contact

apex@mercor.com

Legal disclaimer on the content of worlds

This material is provided for research, educational, and informational purposes only. It consists of hypothetical, simulated financial, legal, and regulatory analyses and illustrative scenarios, including simulated leveraged buyout structures, capital structures, financing terms, valuation ranges, projected returns, potential mergers, acquisitions, divestitures or other strategic transactions, legal memoranda, hypothetical legal advice to a company, and hypothetical correspondence to regulatory agencies. No representation is made that any scenario described here is likely to occur, is being contemplated by any person, or reflects an actual proposed or pending transaction or any legal, regulatory, or compliance risk.

This material does not constitute, and should not be construed as, financial, investment, legal, tax, accounting, or other professional advice. It is not intended to form the basis of an investment decision or contract. The analyses and outputs are based on assumptions, estimates, modeling methodologies, and hypothetical legal scenarios that may prove incorrect. Financial and legal information may be derived from publicly available information and third-party sources that have not been independently verified. Projections, forward-looking statements, scenario outputs, similar financial information, and legal documents, memoranda, and correspondence are hypothetical, inherently uncertain, and provided solely to illustrate how results might change under different assumptions. No representation or warranty, express or implied, is made regarding this material, which is provided on an "as-is" and "as-available" basis.

To the maximum extent permitted by applicable law, Mercor disclaims liability for direct or indirect losses or damages arising from or related to the use of, or reliance on, this material. This includes loss of profits, business, or goodwill, and consequential, incidental, special, punitive, or exemplary damages, even if advised of their possibility. Nothing in this disclaimer limits or excludes liability that cannot be limited or excluded under applicable law.

Robots exclusion statement (human-readable)

To all automated crawlers and bots:

User-Agent: *
Disallow: /

We ask that you:

  • Do not crawl, scrape, index, or download this dataset programmatically.
  • Do not use this dataset for training models or other automated processing without express permission from the dataset owner.
Task
mercor/world228-sm-task08-11f60d76
mercor/world112-1-task05-na-cf44bbb8
mercor/world419-um-04-d0f9f0ee
mercor/cw134-aditi-01-6914a79e
mercor/world425-jcf-01-e1dfe9ed
mercor/world-419-um-01-ddd3e22e
mercor/task-ymtecb81-c206b308
mercor/world127-am-task03-cab85eb0
mercor/world431-jcf-01-3d85cec9
mercor/world224-hs-09-56aefeb6
mercor/world133-ln-04-009ff905
mercor/world132-da-task07-d7bf24d6
mercor/taskworld130-camillemoingeon-5-0d2a857d
mercor/world434-ah-05-0ad4be07
mercor/world132-da-task03-3db0b339
mercor/world-134-nancy-task-02-3f38c565
mercor/world225-km-06-19899d4e
mercor/world227-tg-07-c2f84823
mercor/world-129-cy-task-6-13a7957a
mercor/world-128-sf-task-2-7bab9d72
mercor/world225-av-02-55d8f94f
mercor/world425-jcf-02-9f502d4d
mercor/world423-dpm-01-178fcb61
mercor/world418-tk-01-50b3b6a8
mercor/world431task-el-01-6ec3da62
mercor/world421-ap-02-930732dc
mercor/world219-tg-04-d08a16f0
mercor/world416-js-02-029d3e95
mercor/world434-ah-01-f0eb6d9f
mercor/task-14-24dc5ca2
mercor/world224-hs-11-7bf318b0
mercor/task-w135-camille-moingeon-4-1cf93f93
mercor/world221-oa-2-44073f52
mercor/world226-td-01-eb948f85
mercor/world-127-am-task-02-ec3c5eb0
mercor/task-yhzc9d1a-531a5446
mercor/lawworld433-anb-01-5d12cf8d
mercor/task-awys8050-cd3c370b
mercor/world223-es-05-e6fedd35
mercor/world132-pm-task07-b206dfe1

Displaying 40 of 240 tasks