Benchmarks / MCP Atlas

MCP Atlas (MCP-Atlas)

Built by Scale AI · released 19 Sep 2025

Single-turn requests that an agent can only answer by chaining tools on real Model Context Protocol (MCP) servers: web search, maps, weather, file systems, MongoDB and Airtable databases, Notion, Slack, email, arXiv and PubMed, market data, Git and GitHub, and code runners. Each request needs 3 to 6 tool calls, usually across several servers.

Frontier

88.1%

Pass rate

Muse Spark 1.1

9 Jul 2026 · Source: Scale AI (benchmark maintainers)

Rank 1 on the board seen 2026-09-24, plus or minus 1.95. Fable 5.1 (added 2026-09-04) shares rank 1 at 87.2% plus or minus 2.05; claude-opus-5 (xhigh) is at 85.8%. The page header still says 83.6%.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Scale AI

MCP Atlas: Pass rate over time, 3 recorded results. 0%20%40%60%80%100%Oct 2025Jan 2026Apr 2026Jul 2026 GPT-5: 44.5% (19 Sep 2025) Muse Spark: 82.2% (8 Apr 2026) Muse Spark 1.1: 88.1% (9 Jul 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/mcp-atlas/"><img src="https://canagentswork.com/og/benchmarks-mcp-atlas.png" width="600" height="315" alt="MCP Atlas: the best result is 88.1% (Muse Spark 1.1, 9 Jul 2026)." loading="lazy"></a>

Markdown:

[![MCP Atlas: the best result is 88.1% (Muse Spark 1.1, 9 Jul 2026).](https://canagentswork.com/og/benchmarks-mcp-atlas.png)](https://canagentswork.com/benchmarks/mcp-atlas/)

What it measures

Whether a model can find the right tools in a noisy menu, call them correctly, recover from errors, and combine the outputs into a correct answer. 1,000 human-written tasks (500 public, 500 private) run against 36 real MCP servers in Docker with fixed data, grouped by Scale as search and fetch (brave_search, ddg_search, exa, fetch, weather, google-maps), analytics (mongodb, airtable, calculator), productivity (filesystem, notion, slack, google-workspace, arxiv, pubmed), financial (twelvedata, alchemy), and coding (git, github, mcp-code-executor, cli-mcp-server, e2b-server). Each task exposes 10 to 25 tools, of which 3 to 7 are needed and the rest are distractors.

Pass rate: Share of all 1,000 tasks where the final answer covers at least 75% of the ground-truth claims. An LLM judge scores each claim 0, 0.5, or 1; coverage is the mean, and a task passes at 0.75 or above. A completion rate at the task level, with partial credit inside the threshold. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools
Grading
llm-judge, rubric
Tasks
1,000
Contamination
Half of the tasks are held out. Stateful servers are seeded with fixed data so answers stay stable; search tasks hit live services.
Reuse
The 500 public tasks are on Hugging Face under CC BY 4.0 (prompts, enabled tools, claims, and reference trajectories). The other 500 tasks are private. (open-cc-by)

Limits to keep in mind

  • The page's header ("83.6% Top Pass Rate") and Key Findings ("Top performer: Claude Opus 4.5 with 62.3%") lag the leaderboard table, which shows 88.1% for Muse Spark 1.1. Source
  • In April 2026 Scale changed the judge, added retries for transient tool errors, and replaced the 20-turn limit with a budget of 100 tool calls, then re-scored every model. Scores from before the change are not comparable. Source
  • Grading uses an LLM judge (Gemini-2.5-Pro on the page; the open harness defaults to gemini-3.1-pro-preview) against claim lists. A task can pass with a quarter of its claims wrong. Source
  • Tasks are single-turn lookups and computations with a known answer, not open-ended work products. Coverage is fixed to these 36 servers (maintainers' list): airtable, alchemy, arxiv, brave-search, calculator, cli-mcp-server, clinicaltrialsgov-mcp-server, context7, ddg-search, desktop-commander, e2b-server, exa, fetch, filesystem, git, github, google-maps, google-workspace, lara-translate, mcp-code-executor, mcp-server-code-runner, memory, met-museum, mongodb, national-parks, notion, open-library, osm-mcp-server, oxylabs, pubmed, slack, twelvedata, weather, weather-data, whois, wikipedia. The page counts 220 tools; the maintainers' list counts 307 tools on the same 36 servers. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemPass rateDateSource
Muse Spark 1.1 · frontier88.1%9 Jul 2026Scale AI · primary
Muse Spark82.2%8 Apr 2026Scale AI · primary
GPT-544.5%19 Sep 2025Scale AI · primary

Timeline

  • 9 Jul 2026 — Muse Spark 1.1 passes 88.1% of MCP Atlas tasks. Source
  • Apr 2026 — MCP Atlas re-scored with a new judge and a 100-tool-call budget. Source
  • 19 Sep 2025 — Scale AI launches MCP Atlas; GPT-5 passes 44.5% of real MCP tool-use tasks. Source

Where it sits in the atlas

Work ladder: Context only. It spans many occupations, so it does not set the step of a job family. How the ladder works

Office and administrative supportData and analytics (partial)Software engineering (partial)Finance and accounting (partial)

MCP tool useWeb researchEnterprise workflows

Last checked 24 Sep 2026 against 6 primary sources. See an error? Tell us.

How to cite

Credit the original work first: MCP Atlas by Scale AI (https://arxiv.org/abs/2602.00933).

Then, if you used this page:

Can Agents Work. "MCP Atlas: frontier results and sources." https://canagentswork.com/benchmarks/mcp-atlas/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-mcp-atlas,
  title        = {{MCP Atlas: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/mcp-atlas/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}