Benchmarks / Remote Labor Index
Remote Labor Index (RLI)
Built by CAIS and Scale AI · released Oct 2025
Real freelance projects, collected from professionals on Upwork, that an agent must deliver end to end. Trained human evaluators compare each AI deliverable with the deliverable that a paid professional made.
Frontier
15.8%
Automation rate
Claude Fable 5
1 Jul 2026 · Source: CAIS (benchmark maintainers)
CAIS says the three new models were paired with stronger agent scaffolding.
The best result is below 20% of the ceiling, or below half of human parity.
What it measures
Whether a reasonable client would accept the agent's deliverable, compared with a gold-standard deliverable from a professional freelancer. Projects come from 23 Upwork domains, for example 3D and CAD, architecture, graphic design, video and animation, audio, data analysis, and web apps.
Automation rate: Share of projects where the AI deliverable is judged at least as good as the professional deliverable. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- cli, computer-use
- Grading
- human-expert, pairwise-human
- Tasks
- 240
- Human reference
- A professional freelancer's accepted deliverable. Mean human completion time 28.9 hours (median 11.5 hours). Mean project value $632.60 (median $200).
- Contamination
- Leaderboard scores use a private set of 230 projects.
- Reuse
- Scores come from 230 private projects. 10 public projects and the open-source evaluation platform are released for qualitative analysis. (cite-only)
Limits to keep in mind
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Where sources disagree
CAIS, a co-author of the benchmark, reports 15.8% for Claude Fable 5. Two aggregator pages report 16.1%. We did not find the cause. A later re-grade is one possible cause.
- 15.8 — safe.ai (primary) · CAIS post dated 2026-07-01.
- 16.1 — www.benchleader.com (third-party) · BenchLeader page, seen 2026-09-23.
- 16.1 — ai-intensify.com (third-party) · News article, seen 2026-09-23.
We show 15.8. Status: open.
Timeline
- 1 Jul 2026 — Best Remote Labor Index score rises to 15.8%. Source
Where it sits in the atlas
Design, media, and writingArchitecture and engineering (partial)Data and analytics (partial)Software engineering (partial)
Go to the source
- Website www.remotelabor.ai
- Paper arxiv.org
- Full leaderboard labs.scale.com
- Announcement safe.ai
Last checked 23 Sep 2026 against 2 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: Remote Labor Index by CAIS and Scale AI (https://arxiv.org/abs/2510.26787).
Then, if you used this page:
Can Agents Work. "Remote Labor Index: frontier results and sources." https://canagentswork.com/benchmarks/remote-labor-index/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-remote-labor-index,
title = {{Remote Labor Index: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/remote-labor-index/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}