Method

How we build the atlas, and the rules behind every color and label. If a rule changes, we say so on the timeline and version the change.

1. The territory: jobs

Jobs come from the SOC 2018 system through the O*NET 31.0 Database (997 civilian occupations; military occupations are out of scope because BLS OEWS does not cover them). We group occupations into 17 job families with ordered rules on SOC codes. Area on the atlas is US employment from the BLS OEWS May 2025 national estimates (830 detailed occupations). O*NET sub-occupations share the employment figure of their SOC group, so they change the pages, not the area.

2. Exploration: which benchmarks cover which jobs

Each benchmark maps to one or more job families: primary for its main family and partial for others. Some benchmarks also map to specific occupations. A family with no mapped benchmark shows as unexplored (hatched). That means nobody has measured the work in public in a way we can cite. It does not mean agents cannot do it.

3. Two grains: task and project

4. Benchmark status

Status compares the frontier result with a reference point that each benchmark declares. A ceiling is a hard maximum, such as 100% of tasks. A parity line is human-expert performance, such as 50% wins and ties against experts.

Saturated
The best result is at 90% or more of the ceiling.
Strong
The best result is at 60–90% of the ceiling, or at or above human parity.
Emerging
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
Open
The best result is below 20% of the ceiling, or below half of human parity.
Unrated
No fixed reference point (for example Elo scores or field signals).
Solved
Its maintainers or a major evaluator declared it solved.
Retired
Maintainers or a major user stopped using it as a frontier measure.
No results yet
We have not recorded a result yet.

5. Readings

A reading is our written judgment of what the benchmarks mapped to a family show, at one grain. We write it, date it, and list its evidence. It is not a formula. Rules:

How we pick a level, using only the benchmarks at the reading's grain:

  1. Unexplored: no benchmark at this grain maps to the family.
  2. Not yet: the benchmarks that map most directly to the family show frontier results in the open band, or the best agents lose to a human baseline.
  3. Partial: the evidence is mixed, or only adjacent benchmarks (partial mappings) show success, or success is shown only on narrow parts of the work.
  4. Strong: at least two primary-mapped benchmarks from independent institutions are strong, saturated, or solved, and no primary-mapped benchmark at this grain is open.

Benchmarks with no reference point (Elo scores, field signals, behavior rates) add context but do not set the level by themselves. A stale result still counts until a newer one replaces it. The reading says so when a result is old.

Strong
Agents succeed on most benchmark tasks. Limits remain in reliability, scope, or cost.
Partial
Agents succeed on some scoped tasks. Reliable or end-to-end work is not shown.
Not yet
The best agents fail most tasks on the best available benchmarks.
Unexplored
No benchmark covers this work yet.

6. Sources and provenance

7. What we do not do (yet)

8. Known limits

9. Corrections

If a number, summary, or mapping is wrong, tell us. We aim to correct errors within 48 hours of a report and record the change. Every entry shows the date we last checked it against its sources.