Benchmarks / CRMArena-Pro

CRMArena-Pro

Built by Salesforce AI Research · Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal, et al. · released 24 May 2025

Salesforce AI Research's benchmark for agents that work inside a CRM. Agents answer sales, service, and configure-price-quote requests by querying two synthetic Salesforce orgs (one B2B, one B2C) through the Salesforce API, in single-turn and multi-turn settings with a simulated user.

Frontier

58.3%

Single-turn task success (B2C org)

gemini-2.5-pro (ReAct) · harness: ReAct

24 May 2025 · Source: Salesforce AI Research (benchmark maintainers)

Table 2 of the paper. B2B single-turn 54.1. Multi-turn 35.1 (B2B) and 30.0 (B2C). Workflow execution alone reached 83.0 (B2B) and 90.0 (B2C) single-turn. No newer primary results exist.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

CRMArena-Pro: Single-turn task success (B2C org) over time, 3 recorded results. 0%20%40%60%80%100%May 2025May 2025May 2025May 2025Jun 2025Jun 2025 gpt-4o (ReAct): 29.2% (24 May 2025) o1 (ReAct): 49.5% (24 May 2025) gemini-2.5-pro (ReAct): 58.3% (24 May 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/crmarena-pro/"><img src="https://canagentswork.com/og/benchmarks-crmarena-pro.png" width="600" height="315" alt="CRMArena-Pro: the best result is 58.3% (gemini-2.5-pro (ReAct), 24 May 2025)." loading="lazy"></a>

Markdown:

[![CRMArena-Pro: the best result is 58.3% (gemini-2.5-pro (ReAct), 24 May 2025).](https://canagentswork.com/og/benchmarks-crmarena-pro.png)](https://canagentswork.com/benchmarks/crmarena-pro/)

What it measures

Task completion: whether the agent's final answer matches the gold answer for each query. The 19 task types cover four skills: workflow execution, policy compliance, text understanding, and database querying. A second track checks whether agents refuse to reveal confidential data. The headline is the single-turn success rate on the B2C org, averaged over the four skills. Multi-turn scores, where a persona-driven simulated user holds back details, are much lower.

Single-turn task success (B2C org): Share of single-turn queries on the B2C org where the agent's answer matches the gold answer, averaged over the four business skills. The paper also reports B2B and multi-turn results. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, chat
Grading
automated-tests, llm-judge
Tasks
4,280
Contamination
Queries, gold answers, and the org data are public on Hugging Face. The orgs hold synthetic data that gpt-4o generated from Salesforce schemas.
Reuse
CC BY-NC 4.0 (repository LICENSE.txt). Research use only. (cite-only)

Limits to keep in mind

  • Published scores come from the May 2025 paper (o1, gpt-4o, Gemini 2.5, Llama 3.1 and 4). The Hugging Face leaderboard covers only the original CRMArena, so newer models are not tracked here. Source
  • The org data is synthetic and generated by gpt-4o. 66.7% of CRM experts rated the B2B data as realistic and 62.3% rated the B2C data as realistic. Source
  • Multi-turn runs use an LLM user simulator. A manual check of 20 trajectories found one simulator error (5%). gpt-4o also extracts answers and judges confidentiality refusals. Source
  • Agents run through a ReAct scaffold with API access only. GUI access to the orgs was withdrawn after a Salesforce system update. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemSingle-turn task success (B2C org)DateSource
gpt-4o (ReAct)
harness: ReAct
29.2%24 May 2025Salesforce AI Research · primary
o1 (ReAct)
harness: ReAct
49.5%24 May 2025Salesforce AI Research · primary
gemini-2.5-pro (ReAct) · frontier
harness: ReAct
58.3%24 May 2025Salesforce AI Research · primary

Timeline

  • 24 May 2025 — Salesforce releases CRMArena-Pro. Source

Where it sits in the atlas

Sales and marketingCustomer support (partial)

CRM operationsCustomer service

Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.

How to cite

Credit the original work first: CRMArena-Pro by Salesforce AI Research (https://arxiv.org/abs/2505.18878).

Then, if you used this page:

Can Agents Work. "CRMArena-Pro: frontier results and sources." https://canagentswork.com/benchmarks/crmarena-pro/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-crmarena-pro,
  title        = {{CRMArena-Pro: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/crmarena-pro/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}