Can agents work?

This atlas maps 70 agent benchmarks from 62 institutions onto 997 US occupations. Color shows the largest unit of work that agents complete on direct benchmarks. Stripes show how little of the work those benchmarks measure. Area is US employment.

8%of the work in US jobs is measured by a direct agent benchmark (by O*NET work activities).
34%of US jobs are in job families where no direct benchmark gives a success rate.
70benchmarks, with 252 results that each link to a source.
13benchmarks are saturated, solved, or retired. The frontier moves fast.

The Work Atlas

Color: the work ladder

  1. Most projects · 0% of US jobs
    Agents complete 60% or more of whole projects on two direct benchmarks from independent institutions.
  2. Some projects · 0% of US jobs
    Agents complete 20% or more of whole projects on two direct benchmarks from independent institutions.
  3. Nearly all tasks · 0% of US jobs
    Agents complete 90% or more of the tasks (or projects) on two direct benchmarks from independent institutions.
  4. Most tasks · 11% of US jobs
    Agents complete 60% or more of the tasks (or projects) on two direct benchmarks from independent institutions.
  5. Some tasks · 41% of US jobs
    Agents complete 20% or more of the tasks (or projects) on a direct benchmark.
  6. Not yet · 14% of US jobs
    Direct benchmarks exist, but agents complete less than 20% of their tasks.
  7. Unmeasured · 34% of US jobs
    No direct benchmark gives a success rate.

Texture: how much of the work is measured

  • Most of the work is measured (50% or more)
  • Some of the work is measured (20% to 50%)
  • Little of the work is measured (less than 20%)

Coverage is the share of a job family's work, by O*NET work activities, that direct benchmarks test. How we measure it.

Arrow: the last 12 months

A higher step than on 24 Sep 2025, or measured for the first time since then.

A lower step than on 24 Sep 2025, for example after a benchmark was retired.

Area: US employment, BLS OEWS May 2025 (155M jobs). Color and stripes: the work ladder, computed from the direct benchmarks on each family page, with data as of 24 Sep 2026. Work: O*NET 31.0 occupations and work activities. Method.

Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/"><img src="https://canagentswork.com/og/home.png" width="600" height="315" alt="Can agents work? Direct agent benchmarks measure only 8% of the work in US jobs. The Work Atlas maps 70 benchmarks onto the job market." loading="lazy"></a>

Markdown:

[![Can agents work? Direct agent benchmarks measure only 8% of the work in US jobs. The Work Atlas maps 70 benchmarks onto the job market.](https://canagentswork.com/og/home.png)](https://canagentswork.com/)

Full-size atlas image (PNG, 1600 × 1200, with legend and sources): download.

How to cite

Credit the original work first: the benchmarks behind each job family, listed on each family page and at /sources/.

Then, if you used this page:

Can Agents Work. "The Work Atlas: AI agent benchmarks mapped onto US jobs." https://canagentswork.com/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-work-atlas,
  title        = {{The Work Atlas: AI agent benchmarks mapped onto US jobs}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}

Can agents do my job?

Job families

FamilyUS jobsStepWork measuredLast 12 monthsDirect benchmarks
Food service, cleaning, and personal care21MNot yet9%First measured1
Healthcare18MSome tasks1%No change3
Management and business operations18MSome tasks10%First measured3
Office and administrative support15MMost tasks25%No change5
Sales and marketing14MSome tasks3%No change1
Transportation and material moving14MUnmeasuredless than 1%No change1
Construction, installation, and repair13MUnmeasured0%No change0
Education and social services12MUnmeasured0%No change0
Production and manufacturing8.6MUnmeasured0%No change0
Protective service3.8MUnmeasured0%No change0
Finance and accounting3.1MSome tasks34%No change4
Architecture and engineering2.6MSome tasks3%First measured2
Customer support2.6MSome tasks57%No change1
DevOps, SRE, and IT operations2.4MSome tasks21%No change5
Software engineering2.2MMost tasks47%No change9
Design, media, and writing2.0MUnmeasured0%No change0
Science and research1.5MSome tasksless than 1%No change1
Legal1.3MSome tasks41%First measured3
Data and analytics500KSome tasks8%No change2
Farming, fishing, and forestry434KUnmeasured0%No change0
Security191KNearly all tasks35%Up from Not yet7
ML and research engineering37KSome tasks23%No change3

Latest on the frontier

  1. 24 Sep 2026The atlas moves to the work ladder and coverage (method 2)
  2. 22 Sep 2026Muse Spark 1.2 completes 25.42% of Harvey LAB held-out tasks on the Vals AI board
  3. 22 Sep 2026SWE-Bench Pro V2 launches with 642 validated tasks and a locked offline protocol
  4. 22 Sep 2026SWE-Bench Pro V2 is saturated at launch: Opus 5 with Claude Code resolves 99.4%
  5. 22 Sep 2026Scale AI replaces the SWE-Bench Pro public board with V2; the V1 board moves to Legacy
  6. 21 Sep 2026gNucleus AI releases Parametric CAD Bench V3; GPT-6 Astra with Codex leads at 56.87%

Full timeline · RSS

Recent frontier results

All benchmarks

The atlas stands on their work

Every benchmark here was built by someone else. We summarize, map, and link. Start with their work.

aavetisAfterQueryAI Security InstituteAi2AiderAmazonAndon LabsAnthropicApturaARPA-HArtificial AnalysisBoxCAISCarnegie Mellon UniversityCisco ResearchConcordia UniversityCrosby LegalDARPADexmalEVAL SYSgNucleus AIGoogle CloudGray Swan AIHarveyHKUHugging FaceIBM ResearchIIScLaude InstituteLobeHubMartianMercorMetaMETRmicro1Microsoft ResearchMilaNebiusNUS TRAILOpenAIPeking UniversityPrinceton UniversityQueen's UniversityRampReflectionRenmin University of ChinaSalesforce AI ResearchScale AIServiceNowShortcutSierraSnorkel AIStanford UniversityUC BerkeleyUIUCUniversité de MontréalUniversity of MichiganUniversity of TorontoUniversity of WaterlooVals AIWaymoZapier