Benchmarks / SWE-Lancer
SWE-Lancer
Built by OpenAI · Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes Heidecke · released 17 Feb 2025
Real freelance software jobs from Upwork, all from the Expensify codebase, worth $1 million in actual payouts. Independent contributor (IC) tasks ask the model to fix a bug or build a feature; manager tasks ask it to pick the best of several freelancer proposals. The public split, SWE-Lancer Diamond, is what labs report.
Frontier
80%
IC SWE Diamond pass@1
gpt-5.1-codex-max
18 Nov 2025 · Source: OpenAI (benchmark maintainers)
IC SWE Diamond, pass@1 averaged over 3 runs, on the 2025-07-17 offline dataset. The system card shows the value only in a chart labeled 80%. OpenAI is the maintainer and reports its own model. Later OpenAI system cards do not include SWE-Lancer.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
For IC tasks, the model's patch must pass end-to-end browser tests that professional engineers wrote and reviewed three times. For manager tasks, the choice must match the proposal the real hiring manager picked. The paper reports dollars earned; OpenAI's later system cards report pass@1 on the IC SWE Diamond set, which is the metric here. The 2025-07-17 offline revision keeps 198 IC Diamond tasks and removes internet access during runs.
IC SWE Diamond pass@1: Share of Diamond-set independent contributor tasks whose end-to-end tests pass, one attempt per task. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- repo, cli
- Grading
- automated-tests
- Tasks
- 1,488
- Human reference
- Each task's value is the real payout to the freelancer who did it, from $250 to $32,000 in the full set. IC tasks range from 15-minute bug fixes to multi-week features.
- Contamination
- Tasks come from a public open-source repository (Expensify). OpenAI keeps most tasks private and disables internet during runs.
- Reuse
- Diamond split is public in the openai/preparedness repository (MIT). The remaining tasks are a private holdout. (open-mit)
Limits to keep in mind
- One codebase (the open-source Expensify repository), so results say little about other stacks or domains. OpenAI lists this as a limitation. Source
- OpenAI is both maintainer and the only lab that reports results. No other model developer adopted the benchmark, and OpenAI system cards from December 2025 on no longer include it. Source
- Dataset changed on 2025-07-17: 39 of 237 IC Diamond tasks were dropped and internet access was removed, so results before and after are not comparable. Source
- pass@1 is one sample per task, and OpenAI notes significant variance between runs. Later system cards average three runs. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
Where it sits in the atlas
Go to the source
Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.
How to cite
Credit the original work first: SWE-Lancer by OpenAI (https://arxiv.org/abs/2502.12115).
Then, if you used this page:
Can Agents Work. "SWE-Lancer: frontier results and sources." https://canagentswork.com/benchmarks/swe-lancer/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-swe-lancer,
title = {{SWE-Lancer: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/swe-lancer/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}