Timeline of the frontier

Launches, big jumps, saturation, and retirements. Each entry links to its source. RSS.

2026

  1. 22 Sep 2026
    Method changeScale AI releases SWE-Bench Pro V2 with 642 validated tasks

    V2 drops 89 tasks from the 731-task public set, rewrites 529 problem statements, repairs verifiers, and adds a HARD-51 subset. It runs under a locked protocol: no network during the agent phase and re-grading on a fresh sandbox. The leaderboard still shows V1 results. Source

  2. 8 Sep 2026
    Method changeAPEX-Agents 1.1 stops rewarding hedged answers

    Mercor found models raising scores by giving several answers at once ("scattergunning"). Version 1.1 trims the set to 240 audited tasks, adds a judge that scores hedged rubric items as zero, and fixes tools. Scores rise in aggregate and are not comparable with 1.0. Claude Fable 5.1 leads at 68.6% Pass@1. Source

  3. 7 Sep 2026
    Big jumpCyberGym leader reaches 98.5% reproduction rate

    The Creation multi-model agent tops the official leaderboard at 98.47%, up from under 20% at launch in June 2025 and 83.1% for Claude Mythos Preview in April 2026. Several team submissions now exceed 95%, and the maintainers warn that small differences may not be meaningful. Source

  4. 7 Sep 2026
    Big jumpGPT-6 Astra takes the Vending-Bench 2 lead at about $15,500

    Andon Labs reports that GPT-6 Astra finished the simulated year with an average balance of $15,515, the first OpenAI model to lead the board and the largest gap over second place so far (Claude Opus 5, $11,181.87). Andon Labs credits steady negotiation all year and never prepaying suppliers that had gone out of business. Source

  5. Sep 2026
    Method changeGDPval-AA v2.1 re-anchors the Elo scale

    Artificial Analysis pins DeepSeek V4.1 Flash (max) at 1600 Elo and fits ratings with a Crowd-BT model. Rank order is largely unchanged, but v2.1 scores are not comparable with v2 scores, which were anchored to human expert performance at 1000. Source

  6. 25 Jul 2026
    Big jumpFirst verified OSWorld score above 90%

    The Intelligence-Indeed Agent completes 90.19% of OSWorld-Verified tasks at 100 steps in a run verified by the maintainers. A week later Claude Fable 5 posts 85.96% as the best general model. Both are well above the 72.36% human figure from the original study. Source

  7. 8 Jul 2026
    Method changeOpenAI audit estimates about 30% of SWE-Bench Pro tasks are broken

    OpenAI reviews the 731-task public split with an agent pipeline and five human engineers per flagged task. The pipeline flags 200 tasks (27.4%) and humans flag 249 (34.1%) as broken. OpenAI withdraws its earlier recommendation to adopt SWE-Bench Pro. Source

  8. 1 Jul 2026
    Big jumpBest Remote Labor Index score rises to 15.8%

    CAIS reports that Claude Fable 5 completes 15.8% of real freelance projects at a client-acceptable standard. At launch in October 2025, the best agent completed 2.5%. Source

  9. Jul 2026
    SaturatedBrowseComp scores pass 90%

    OpenAI reports 90.4% for GPT-5.6 Sol and 92.2% for its four-agent "ultra" mode. By September 2026 GPT-6 Astra (91.5%) and Claude Opus 5 (90.8%) also sit above 90%, leaving little headroom. Source

  10. 26 Jun 2026
    MilestoneOSWorld 2.0 succeeds OSWorld-Verified

    The XLANG team releases OSWorld 2.0, 108 long-horizon computer-use workflows that take skilled users a median of about 1.6 hours. The best agent, Claude Opus 4.8, completes only 20.6% of tasks at 500 steps, showing how much headroom the older benchmark no longer captures. Source

  11. 23 Jun 2026
    Big jumpSpreadsheetBench V1 leader reaches 83.11%

    Kingsoft Office's Qingqiu Agent scores 83.11% on the full 912-question set, a verified entry. The previous verified leader, Gemini in Google Sheets, scored 70.48% in March 2026. Source

  12. 27 May 2026
    LaunchArtificial Analysis and IBM launch ITBench-AA

    59 Kubernetes incident-diagnosis tasks run in a fixed harness. Every frontier model scores below 50%. Claude Opus 4.7 leads at 47%, followed by GPT-5.5 at 46%. Source

  13. 8 May 2026
    Big jumpMETR reports a 17-hour time horizon and warns the suite is near its ceiling

    METR adds Claude Mythos Preview (early) with a 50% horizon of 1,044.8 minutes (about 17.4 hours, CI 508.9 to 3,304.3) and posts a notice that measurements above 16 hours are unreliable with the current task suite. Claude Opus 4.6 had reached 718.8 minutes in February. Source

  14. 8 May 2026
    LaunchSREGym paper released with 90 live SRE problems

    UIUC and University of Toronto publish SREGym. Frontier agents reach 27% to 61% end-to-end success on the 90-problem suite. On failures unique to SREGym (hardware, metastable, concurrent), the best agent falls to 28%. Source

  15. 6 May 2026
    RetiredTerminal-Bench 2.1 fixes 28 tasks and replaces 2.0

    The maintainers release 2.1 after finding defects in 28 of the 89 tasks in 2.0: changed external dependencies, resource budgets too small for valid solutions, and instructions that did not match tests. Most agent-model pairs score higher on 2.1; Claude Code with Opus 4.6 gains 12.1 points. Terminal-Bench 3.0 (July 2026) and 4.0 (August 2026) followed. Source

  16. 24 Apr 2026
    Method changeMLE-bench pauses new leaderboard submissions

    OpenAI stops taking new MLE-bench leaderboard submissions while it builds a process to make sure submissions are fair and comparable. A v2 release with batched fixes is planned in the frontier-evals repo. Source

  17. 7 Apr 2026
    SaturatedClaude Mythos Preview reaches 100% on Cybench subset

    Anthropic's system card reports 100% pass@1 on its 35-task Cybench subset with 10 trials per task, and says that, given the saturation of the benchmark, it is no longer sufficiently informative. Source

  18. Apr 2026
    Big jumpGPT-5.5 reaches 84.9% wins plus ties on GDPval

    In the GPT-5.5 launch post OpenAI reports 84.9% wins plus ties against industry experts, with GPT-5.4 at 83.0% and Claude Opus 4.7 at 80.3%. The post does not say which grader produced the numbers. Source

  19. 11 Mar 2026
    SaturatedGAIA leaderboard passes the 92% human score

    Alibaba Cloud's OPS-Agentic-Search posts 92.36% on the GAIA test set, the first entry above the 92% human figure from the paper. By September 2026 several self-submitted agents sit between 93% and 94%, so the benchmark no longer separates the best systems. Source

  20. 1 Mar 2026
    Big jumpSpider 2.0-Snow leader reaches 96.70%

    Genloop's Sentinel Agent v2 Pro tops the Spider 2.0-Snow table at 96.70%. Fifteen months earlier the maintainers' Spider-Agent with o1-preview scored 23.58%. The Snow variant now has little headroom. Source

  21. 26 Feb 2026
    LaunchMartian launches Code Review Bench

    Martian publishes an open benchmark for AI code review tools with an offline set of 50 PRs and 173 golden comments, plus an online set built from fresh GitHub PRs. The repo, judge prompts, and results are MIT licensed. Source

  22. 23 Feb 2026
    Big jumpMLE-bench leader earns medals in 64.44% of competitions

    Baidu's Famou-Agent 2.0 with Gemini-3-Pro-Preview tops the MLE-bench table at 64.44% any-medal rate. Sixteen months earlier the best agent scored 17.12%. Source

  23. 23 Feb 2026
    RetiredOpenAI stops reporting SWE-bench Verified

    OpenAI says SWE-bench Verified no longer measures frontier coding progress. Its audit found flawed tests in 59.4% of 138 hard tasks, and probes showed frontier models from three labs reproducing gold patches from training data. OpenAI asks other developers to stop reporting it and points to SWE-Bench Pro instead. Source

  24. Feb 2026
    Method changeSierra fixes 50+ airline and retail tasks for τ³-bench

    Sierra audited the airline and retail domains with the SABER (τ-Bench Verified) team at Amazon and fixed 27 airline and 26 retail tasks with wrong expected actions, ambiguous instructions, or impossible constraints. Airline pass^1 rose by 14 to 20 points for the three re-run models, so scores before and after are not comparable. Telecom was not changed. Source

  25. 29 Jan 2026
    Method changeMETR releases Time Horizon 1.1 with 228 tasks

    METR grows the task suite from 170 to 228 tasks, doubles the number of 8-hour-plus tasks to 31, and moves from Vivaria to Inspect. Estimates shift: the post-2023 doubling time drops from 165 to 131 days, and Claude Opus 4.5 moves from 289 to 320 minutes. Source

  26. 21 Jan 2026
    LaunchMercor releases APEX-Agents

    Mercor, with Box and Harvey, releases 480 long-horizon tasks from investment banking, consulting, and corporate law inside simulated project workspaces. The best agent, Gemini 3 Flash, passes 24.0% of tasks on one attempt. Source

2025

  1. 3 Dec 2025
    SolvedHAL declares CORE-Bench solved

    Claude Opus 4.5 in a Claude Code scaffold scores 77.78% on CORE-Bench-Hard, nearly double its 42.22% with HAL's CORE-Agent. After HAL fixed grading errors in 8 tasks and removed 1 broken task, the score is 95.5%. HAL treats the benchmark as solved and plans to open a private test set. Source

  2. Dec 2025
    LaunchArtificial Analysis launches GDPval-AA

    Artificial Analysis opens an independent Elo leaderboard that runs OpenAI's public GDPval tasks through its open-source Stirrup agent harness and grades deliverables with blind pairwise LLM judging. Claude Opus 4.5 leads at launch. Source

  3. Dec 2025
    MilestoneGPT-5.2 Thinking passes the 50% line on GDPval

    OpenAI reports that GPT-5.2 Thinking beats or ties industry professionals in 70.9% of GDPval comparisons, judged by expert humans. It is the first model OpenAI reports above the 50% parity line. Source

  4. 18 Nov 2025
    Big jumpGPT-5.1-Codex-Max reaches 80% on SWE-Lancer IC Diamond

    OpenAI's GPT-5.1-Codex-Max system card reports about 80% pass@1 on IC SWE Diamond tasks, averaged over three runs, up from 55% for GPT-5 and 67% for GPT-5-Codex in the same chart. This is the last OpenAI system card to report SWE-Lancer. Source

  5. 18 Nov 2025
    LaunchAndon Labs releases Vending-Bench 2

    The second Vending-Bench adds adversarial suppliers, negotiation, delivery delays, supplier bankruptcies, and refund demands, and scores the bank balance after one simulated year. Gemini 3 Pro led at launch with $5,478.16, ahead of Claude Sonnet 4.5 ($3,838.74) and Grok 4 ($1,999.46). Source

  6. 7 Nov 2025
    LaunchTerminal-Bench 2.0 and Harbor released

    The Terminal-Bench team releases 2.0, a harder and more carefully verified set of 89 terminal tasks, together with Harbor, a new package for running agents in cloud containers. The best verified row for a pre-launch model is Codex CLI with GPT-5 at 49.6%. Source

  7. 2 Nov 2025
    LaunchCodeClash launches goal-oriented coding tournaments

    Stanford and Princeton researchers release CodeClash, where models evolve codebases over 15-round tournaments in six game arenas. Across 1,680 tournaments, Claude Sonnet 4.5 leads with the highest Elo, followed by GPT-5 and o3. Top models lose every round to expert human bots. Source

  8. 30 Oct 2025
    Method changeCVE-Bench v2.0 closes grading shortcuts

    The maintainers hardened the outbound-request check and the SQL-injection check after finding that agents could pass without real exploits. GPT-4o agent success rates fell by up to 32.5 points after the fixes. Source

  9. 23 Oct 2025
    LaunchImpossibleBench measures how often coding agents game their tests

    CMU and Anthropic researchers release tasks where tests contradict the specification, so any pass is a shortcut. GPT-5 passes 54.0% of the conflicting SWE-bench tasks by cheating. Stricter prompts cut cheating sharply. Source

  10. 2 Oct 2025
    Big jumpGPT-5 reaches 95.8% telecom pass^1 on τ²-bench

    Sierra's own run of GPT-5 (evaluated 2025-08-09, added to the leaderboard 2025-10-02) scored 95.8% pass^1 and 85.1% pass^4 on telecom, up from about 50% pass^1 four months earlier. Telecom pass^1 has stayed near 98% for the best models since then. Source

  11. 25 Sep 2025
    LaunchOpenAI releases GDPval

    OpenAI publishes GDPval, 1,320 real work tasks from 44 occupations in 9 US industries, with a public 220-task gold subset. Blinded experts rate Claude Opus 4.1 better than or equal to the professional's deliverable in 47.6% of comparisons. Source

  12. 21 Sep 2025
    LaunchScale AI releases SWE-Bench Pro

    Scale AI publishes SWE-Bench Pro with 1,865 long-horizon tasks from 41 repositories, including a 731-task public set from GPL repositories and a 276-task private set from startups. Under a 50-turn, $2 cap, the best public-set result is 23.3% (GPT-5, medium reasoning). Source

  13. 28 Jul 2025
    Method changeOSWorld becomes OSWorld-Verified

    The HKU XLANG team fixes about 300 reported problems in OSWorld tasks and checkers, moves evaluation to a parallel AWS setup, and re-runs all baselines. The report names CoACT-1 as the best agent at 60.76%, about 84% of the roughly 72% human figure. Source

  14. 20 Jul 2025
    LaunchAIDev dataset of agent pull requests released

    Queen's University researchers publish AIDev, 456,535 pull requests by five coding agents across 61,453 GitHub repositories. In popular repositories, agent PRs are merged less often than human PRs (Codex 65.3% versus 76.8% for humans). Source

  15. 9 Jun 2025
    LaunchSierra releases τ²-bench with a dual-control telecom domain

    τ²-bench adds a telecom customer-service domain where both the agent and the simulated user hold tools. At launch the best telecom pass^1 was about 50% (o4-mini, gpt-4.1-mini, and Claude 3.7 Sonnet at 49%), and GPT-4.1 dropped from 74% on retail to 34% on telecom. Source

  16. 3 Jun 2025
    LaunchUC Berkeley releases CyberGym with 1,507 real vulnerabilities

    Agents must write proof-of-concept inputs that reproduce real OSS-Fuzz vulnerabilities in 188 projects. The best agent-model pairs at launch reproduce fewer than 20% of them. Source

  17. 26 May 2025
    LaunchPR Arena starts tracking agent pull requests on GitHub

    The PRarena repo records its first data point. At that time Codex had 51,548 ready PRs with 85.8% merged, and Copilot had 2,099 ready PRs with 75.94% merged. Source

  18. 26 May 2025
    LaunchNebius launches SWE-rebench with fresh GitHub issues

    Nebius publishes an automated pipeline that collects new issue and pull request pairs from Python repositories, plus a leaderboard that runs every model in the same ReAct scaffold five times. In the paper, GPT-4.1 leads the January 2025 window at 31.1%. Source

  19. 24 May 2025
    LaunchSalesforce releases CRMArena-Pro

    CRMArena-Pro expands CRMArena to 19 sales, service, and CPQ task types across B2B and B2C Salesforce orgs, adds multi-turn users and confidentiality checks. The best agent (gemini-2.5-pro with ReAct) completed 58.3% of single-turn B2C queries and about 35% in multi-turn. All models showed near-zero confidentiality awareness without special prompting. Source

  20. 20 May 2025
    LaunchVals AI releases the Finance Agent Benchmark

    Vals AI publishes 537 expert-written financial research questions and an agent harness with EDGAR and web search tools. The best model, OpenAI o3, answered 46.8% correctly at an average cost of $3.79 per query. Source

  21. 2 Apr 2025
    LaunchOpenAI releases PaperBench for replicating ICML papers

    Agents must replicate 20 ICML 2024 papers from scratch. Claude 3.5 Sonnet with a basic scaffold scores 21.0%, and o1 with a 36-hour limit scores 26.0%. ML PhDs reached 41.4% on a 3-paper subset after 48 hours. Source

  22. Apr 2025
    LaunchOpenAI releases BrowseComp

    OpenAI open-sources 1,266 hard-to-find, easy-to-verify browsing questions. Deep research answers 51.5%, o1 9.9%, and GPT-4o with browsing 1.9%. Human trainers solved 29.2% within two hours. Source

  23. 31 Mar 2025
    LaunchUIUC releases CVE-Bench with 40 critical web CVEs

    Agents must exploit real critical-severity vulnerabilities in sandboxed web applications. The best GPT-4o agent framework exploits 12.5% of CVEs with five attempts when given a vulnerability description. Source

  24. 18 Mar 2025
    LaunchMETR introduces the 50% task-completion time horizon

    METR's paper "Measuring AI Ability to Complete Long Software Tasks" defines the time horizon metric. Claude 3.7 Sonnet had a 50% horizon of around 50 minutes, and the frontier horizon had doubled about every seven months since 2019. Source

  25. 18 Feb 2025
    LaunchAmbig-SWE tests whether coding agents ask for clarification

    CMU researchers release an underspecified variant of SWE-Bench Verified with a simulated user. Models rarely ask questions unless prompted, and most cannot tell a vague issue from a complete one. When they do interact, resolve rates rise sharply. Source

  26. 17 Feb 2025
    LaunchOpenAI releases SWE-Lancer with $1 million of Upwork tasks

    OpenAI publishes 1,488 real freelance software tasks worth $1 million in actual payouts, with a public Diamond split worth $500,800. The best model, Claude 3.5 Sonnet, solves 26.2% of Diamond IC tasks and earns $208,050 across the Diamond set. Source

  27. 7 Feb 2025
    LaunchIBM Research releases ITBench

    IBM Research and UIUC publish ITBench, a framework of live Kubernetes scenarios for SRE, compliance (CISO), and FinOps agents. The first paper reports that agents resolve 13.8% of SRE scenarios. Source

  28. 24 Jan 2025
    LaunchStanford releases MedAgentBench

    MedAgentBench gives agents 300 physician-written tasks in a FHIR-compliant virtual EHR with 100 de-identified patients. Claude 3.5 Sonnet v2 led with a 69.67% success rate; models did much better on record lookups (up to 85%) than on tasks that change records. Source

  29. 12 Jan 2025
    LaunchAIOpsLab paper released with 48 cloud operations problems

    Microsoft Research and partners publish AIOpsLab, a framework that deploys microservices, injects faults, and evaluates agents on detection, localization, root cause analysis, and mitigation. The best agent in the paper reaches 59.32% accuracy. Source

2024

  1. 21 Dec 2024
    LaunchAider launches the polyglot leaderboard

    Paul Gauthier replaces aider's saturating Python-only benchmark with 225 hard Exercism exercises in six languages. o1 with high reasoning effort leads at 61.7%, ahead of Claude 3.5 Sonnet at 45.3%. Source

  2. 18 Dec 2024
    LaunchCMU releases TheAgentCompany

    CMU publishes a simulated software company with GitLab, Plane, ownCloud, RocketChat, and language-model coworkers, plus 175 work tasks. The best agent, OpenHands with Claude 3.5 Sonnet, completes 24.0% of tasks. Source

  3. 22 Nov 2024
    LaunchMETR releases RE-Bench with 71 human expert baselines

    Seven ML research engineering environments with matched human and agent conditions. Agents score about 4 times the human average at a 2-hour budget, but humans narrowly pass the best agent at 8 hours and reach about twice the agent score at 32 hours. Source

  4. 12 Nov 2024
    LaunchSpider 2.0 launches with 632 enterprise text-to-SQL problems

    HKU's XLANG Lab and partners release Spider 2.0. The paper reports that its code agent with o1-preview solves 21.3% of tasks, against 91.2% on Spider 1.0 and 73.0% on BIRD. Source

  5. 9 Oct 2024
    LaunchOpenAI releases MLE-bench with 75 Kaggle competitions

    The best setup at launch, o1-preview with the AIDE scaffold, reaches a medal in about 17% of competitions (16.9% in the paper, 17.12% in the repo table). Source

  6. 17 Sep 2024
    LaunchPrinceton releases CORE-Bench for computational reproducibility

    270 tasks from 90 published papers at three difficulty levels. The best agent, CORE-Agent with GPT-4o, reaches 21.48% on the hardest level. Source

  7. 15 Aug 2024
    LaunchStanford releases Cybench with 40 professional CTF tasks

    The best agent (Claude 3.5 Sonnet) solves 17.5% of tasks unguided; GPT-4o solves 12.5%. Agents solve tasks that took human teams up to 11 minutes; the hardest task took humans 24 hours 54 minutes. Source

  8. 13 Aug 2024
    LaunchOpenAI and the SWE-bench team release SWE-bench Verified

    93 developers screen 1,699 SWE-bench tasks. The 500 tasks that pass become SWE-bench Verified. GPT-4o resolves 33.2% with the best open-source scaffold, double its score on the original SWE-bench. Source

  9. 21 Jun 2024
    LaunchSpreadsheetBench launches with 912 real spreadsheet questions

    Renmin University and partners release SpreadsheetBench. In the paper GPT-4o scores 18.35% (soft) and 15.02% (hard) overall, and Copilot in Excel about 20% on a subset, while Excel experts score 71.33% and 62.00% on a 50-question subset. Source

  10. 12 Mar 2024
    LaunchServiceNow Research releases WorkArena and BrowserGym

    ServiceNow Research publishes WorkArena, 33 browser tasks on a live ServiceNow instance, together with the BrowserGym environment for web agents. WorkArena++ (682 compositional tasks) follows in July 2024. Source

2023

  1. 21 Nov 2023
    LaunchMeta and Hugging Face release GAIA

    GAIA offers 466 questions that need browsing, file reading, and tool use. Humans score 92% while GPT-4 with plugins scores 15%. A leaderboard with private test answers opens on Hugging Face. Source