Benchmarks / BrowseComp
BrowseComp
Built by OpenAI · Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, et al. · released Apr 2025
1,266 hard fact-finding questions with short, single answers. Each answer is hard to find but easy to verify, so a browsing agent must search persistently and combine clues from many sites.
Frontier
92.2%
Accuracy
GPT-5.6 Sol Ultra · harness: ultra (four parallel agents)
Jul 2026 · Source: OpenAI (benchmark maintainers)
GPT-5.6 launch post. "Ultra" coordinates four agents in parallel by default. GPT-5.6 Sol without Ultra scored 90.4% in the same table. OpenAI is both the benchmark maintainer and the model developer.
The best result is at 90% or more of the ceiling.
What it measures
Share of questions answered correctly. Human trainers wrote "inverted" questions from a seed fact and several constraints (for example, find a paper by the authors' universities and venue). Every question was checked to be unsolvable by GPT-4o, o1, and an early deep research model at the time, and not findable on the first page of five simple searches. Grading compares the short answer with the reference.
Accuracy: Share of the 1,266 questions answered correctly. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- browser
- Grading
- automated-tests
- Tasks
- 1,266
- Human reference
- Trainers who did not write the question solved 29.2% of problems within a two-hour limit and without AI help. Of solved problems, 86.4% matched the reference answer.
- Contamination
- Questions and answers are public in the simple-evals repository, so later models may have trained on them. OpenAI has not published a contamination analysis.
- Reuse
- Released in OpenAI's simple-evals repository, which lists BrowseComp under the MIT License. (open-mit)
Limits to keep in mind
- Answers are short strings, so the benchmark does not test long answers, ambiguity, or real user queries. OpenAI calls it an incomplete but useful measure of browsing. Source
- OpenAI builds the benchmark and reports its own models. Later scores come from OpenAI launch posts, and the top score uses four agents running in parallel. Source
- The best scores are now above 90%, so the benchmark is close to saturation. Source
- Deep Research, the launch leader, was trained on data that targets BrowseComp-style tasks. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
Where it sits in the atlas
Go to the source
Last checked 23 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: BrowseComp by OpenAI (https://arxiv.org/abs/2504.12516).
Then, if you used this page:
Can Agents Work. "BrowseComp: frontier results and sources." https://canagentswork.com/benchmarks/browsecomp/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-browsecomp,
title = {{BrowseComp: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/browsecomp/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}