Benchmarks / MCPMark Verified
MCPMark Verified (MCPMark)
Built by EVAL SYS, LobeHub and NUS TRAIL · Zijian Wu, Xiangyan Liu, Xinyuan Zhang, Lingjun Chen, et al. · released 12 Jun 2026
Multi-step work on five real MCP servers: edit Notion pages and databases, run GitHub repository and CI workflows, administer and query a PostgreSQL database, browse and act on websites through Playwright, and organize files on a file system. Each task starts from a prepared state, and a script checks the end state.
Frontier
96.1%
Pass@1 (single run)
kimi-k3-max · harness: MCPMarkAgent tool-calling loop
20 Jul 2026 · Source: EVAL SYS (benchmark maintainers)
Rank 1 on the Verified board updated 2026-07-20 (122 of 127 tasks). Per service: Filesystem 93.33, GitHub 95.65, Notion 92.86, Playwright 100.00, Postgres 100.00. Single run, no interval. The board lists it as a closed model.
The best result is at 90% or more of the ceiling.
What it measures
Whether an agent can carry out create, read, update, and delete work through MCP tools and leave the system in the required state. 127 tasks: Filesystem 30, Notion 28, Playwright 25, GitHub 23, Postgres 21. Examples: set up an ESLint workflow on all pull requests, write a PostgreSQL function for inventory transfers with audit logging, recolor elements on a Notion page, extract contact details from mixed file formats. The Verified set (June 2026) is a subset of the standard tasks that pins every server version (server-filesystem 2025.12.18, github-mcp-server v0.15.0, notion-mcp-server 1.9.1, playwright/mcp 0.0.68, postgres-mcp 0.3.0) and stabilizes every verifier.
Pass@1 (single run): Share of the 127 tasks whose verification script passes on one run (run-1). A completion rate. The Legacy board also reported Pass@4 and Pass^4 over four runs; the Verified board is single-run. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools, browser, repo
- Grading
- state-check, automated-tests
- Tasks
- 127
- Contamination
- All tasks and verifiers are public on GitHub. GitHub tasks recreate real open-source repositories from state templates.
- Reuse
- Tasks, initial states, and verifiers ship in the GitHub repository under Apache License 2.0. (open-apache)
Limits to keep in mind
- The top three entries exceed 90% on the Verified set, so it separates the best systems little. Source
- Single-run scores with no confidence interval. The board notes that kimi-k2-7-code's filesystem score comes from its second recorded run, and that claude-opus-4-8-max and kimi-k2-6 are scores-only entries without per-task results. Source
- Results before the Verified release (June 2026) used unpinned server versions and older verifiers; the maintainers call them deprecated and not comparable. Source
- Only five MCP servers, and the tasks were written by the benchmark's own team with AI help rather than sampled from real workloads. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
Where it sits in the atlas
Work ladder: Context only. It spans many occupations, so it does not set the step of a job family. How the ladder works
Office and administrative supportSoftware engineering (partial)Data and analytics (partial)
Go to the source
- Website mcpmark.ai
- Paper arxiv.org
- Full leaderboard mcpmark.ai
- Code github.com
Last checked 24 Sep 2026 against 8 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: MCPMark Verified by EVAL SYS, LobeHub, and NUS TRAIL (https://arxiv.org/abs/2509.24002).
Then, if you used this page:
Can Agents Work. "MCPMark Verified: frontier results and sources." https://canagentswork.com/benchmarks/mcpmark-verified/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-mcpmark-verified,
title = {{MCPMark Verified: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/mcpmark-verified/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}