Benchmarks / BrowseComp

BrowseComp

Built by OpenAI · Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, et al. · released Apr 2025

1,266 hard fact-finding questions with short, single answers. Each answer is hard to find but easy to verify, so a browsing agent must search persistently and combine clues from many sites.

Frontier

92.2%

Accuracy

GPT-5.6 Sol Ultra · harness: ultra (four parallel agents)

Jul 2026 · Source: OpenAI (benchmark maintainers)

GPT-5.6 launch post. "Ultra" coordinates four agents in parallel by default. GPT-5.6 Sol without Ultra scored 90.4% in the same table. OpenAI is both the benchmark maintainer and the model developer.

Saturated

The best result is at 90% or more of the ceiling.

BrowseComp: Accuracy over time, 4 recorded results. 0%20%40%60%80%100%Apr 2025Jul 2025Oct 2025Jan 2026Apr 2026Jul 2026Oct 2026 Deep research: 51.5% (Apr 2025) GPT-5.5 Pro: 90.1% (Apr 2026) GPT-5.6 Sol Ultra: 92.2% (Jul 2026) GPT-6 Astra: 91.5% (Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/browsecomp/"><img src="https://canagentswork.com/og/benchmarks-browsecomp.png" width="600" height="315" alt="BrowseComp: the best result is 92.2% (GPT-5.6 Sol Ultra, Jul 2026)." loading="lazy"></a>

Markdown:

[![BrowseComp: the best result is 92.2% (GPT-5.6 Sol Ultra, Jul 2026).](https://canagentswork.com/og/benchmarks-browsecomp.png)](https://canagentswork.com/benchmarks/browsecomp/)

What it measures

Share of questions answered correctly. Human trainers wrote "inverted" questions from a seed fact and several constraints (for example, find a paper by the authors' universities and venue). Every question was checked to be unsolvable by GPT-4o, o1, and an early deep research model at the time, and not findable on the first page of five simple searches. Grading compares the short answer with the reference.

Accuracy: Share of the 1,266 questions answered correctly. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
browser
Grading
automated-tests
Tasks
1,266
Human reference
Trainers who did not write the question solved 29.2% of problems within a two-hour limit and without AI help. Of solved problems, 86.4% matched the reference answer.
Contamination
Questions and answers are public in the simple-evals repository, so later models may have trained on them. OpenAI has not published a contamination analysis.
Reuse
Released in OpenAI's simple-evals repository, which lists BrowseComp under the MIT License. (open-mit)

Limits to keep in mind

  • Answers are short strings, so the benchmark does not test long answers, ambiguity, or real user queries. OpenAI calls it an incomplete but useful measure of browsing. Source
  • OpenAI builds the benchmark and reports its own models. Later scores come from OpenAI launch posts, and the top score uses four agents running in parallel. Source
  • The best scores are now above 90%, so the benchmark is close to saturation. Source
  • Deep Research, the launch leader, was trained on data that targets BrowseComp-style tasks. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemAccuracyDateSource
GPT-6 Astra91.5%Sep 2026OpenAI · primary
GPT-5.6 Sol Ultra · frontier
harness: ultra (four parallel agents)
92.2%Jul 2026OpenAI · primary
GPT-5.5 Pro90.1%Apr 2026OpenAI · primary
Deep research51.5%Apr 2025OpenAI · primary

Timeline

  • Jul 2026 — BrowseComp scores pass 90%. Source
  • Apr 2025 — OpenAI releases BrowseComp. Source

Where it sits in the atlas

Office and administrative support

Web research

Last checked 23 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: BrowseComp by OpenAI (https://arxiv.org/abs/2504.12516).

Then, if you used this page:

Can Agents Work. "BrowseComp: frontier results and sources." https://canagentswork.com/benchmarks/browsecomp/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-browsecomp,
  title        = {{BrowseComp: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/browsecomp/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}