Benchmarks / HWE-Bench

HWE-Bench

Built by Peking University · Fan Cui, Hongyuan Hou, Zizhang Luo, Chenyun Yin, et al. · released 16 Apr 2026

417 hardware bug-repair tasks from Peking University, built from real bug-fix pull requests in six open-source chip projects (OpenTitan, Ibex, CVA6, Caliptra, XiangShan, Rocket Chip) written in Verilog, SystemVerilog, and Chisel. A coding agent gets a bug report inside a containerized copy of the repository and must fix the RTL so that the project's own simulation tests pass.

Frontier

79.9%

Resolved rate

GPT-5.6 Sol max (Codex CLI) · harness: Codex CLI

31 Jul 2026 · Source: Peking University (benchmark maintainers)

333 of 417 tasks resolved. The JSON stores the count; we computed the percent (79.856...) and rounded to one decimal as the site does. Precision 0.832; cost 1178.73 (unit not stated in the JSON; the README calls it an estimated token-equivalent cost). Added on 2026-07-31 with the leaderboard launch, per the README news. OpenTitan 184/245, Ibex 32/35, CVA6 33/35, Caliptra 15/16, XiangShan 42/54, Rocket Chip 27/32.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Peking University

HWE-Bench: Resolved rate over time, 3 recorded results. 0%20%40%60%80%100%Apr 2026May 2026Jun 2026Jul 2026Aug 2026 GPT-5.4 xhigh (Codex CLI): 70.7% (16 Apr 2026) Claude Opus 4.7 max (Claude Code): 74.6% (30 Apr 2026) GPT-5.6 Sol max (Codex CLI): 79.9% (31 Jul 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/hwe-bench/"><img src="https://canagentswork.com/og/benchmarks-hwe-bench.png" width="600" height="315" alt="HWE-Bench: the best result is 79.9% (GPT-5.6 Sol max (Codex CLI), 31 Jul 2026)." loading="lazy"></a>

Markdown:

[![HWE-Bench: the best result is 79.9% (GPT-5.6 Sol max (Codex CLI), 31 Jul 2026).](https://canagentswork.com/og/benchmarks-hwe-bench.png)](https://canagentswork.com/benchmarks/hwe-bench/)

What it measures

Resolved rate: the share of the 417 tasks where the agent's patch makes the failing simulation test pass (fail-to-pass) under the project's native flow (Verilator, or Synopsys VCS for OpenTitan). The agent receives a generated problem statement, not the original issue text, to avoid leaking the fix. The maintainers run each model with its usual agent CLI (Codex CLI, Claude Code, Kimi CLI, or OpenHands) with web search disabled, one attempt per task. The leaderboard also reports file-level precision of the patch and an estimated cost.

Resolved rate: Share of the 417 tasks resolved (the included simulation test passes after the agent's patch). Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli, repo
Grading
automated-tests
Tasks
417
Contamination
Tasks come from public pull-request history, so the fixes exist online. The maintainers give the agent a rewritten problem statement and disable web search in the agent CLIs. The README is not consistent on network access: the 2026-07-31 news says generated tasks "now explicitly prohibit external access", but the setup section says they "allow public network access during agent installation and execution" and keep only the verifier offline.
Reuse
The task dataset on Hugging Face (henryen/hwe-bench) is tagged Apache 2.0. The code is Apache 2.0. Docker images for five projects are public; OpenTitan images are not distributed because that flow needs a Synopsys VCS license. (open-apache)

Limits to keep in mind

  • OpenTitan supplies 245 of the 417 tasks and needs a commercial Synopsys VCS license. Without it, only the 172 Verilator-based tasks can be run. Source
  • One run per row and no confidence intervals on the leaderboard. Rows mix the model and its agent CLI, so they are not clean model rankings. Source
  • The 2026-07-31 update changed the run setup (new Harbor snapshot, no external access) and re-evaluated GPT-5.5, whose score moved from 79.4% to 76.5%. The README names no re-run for the other rows. Source
  • The maintainers call their rows "reference scores", and the site describes no way to submit results. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemResolved rateDateSource
GPT-5.6 Sol max (Codex CLI) · frontier
harness: Codex CLI
79.9%31 Jul 2026Peking University · primary
Claude Opus 4.7 max (Claude Code)
harness: Claude Code
74.6%30 Apr 2026Peking University · primary
GPT-5.4 xhigh (Codex CLI)
harness: Codex CLI
70.7%16 Apr 2026Peking University · primary

Timeline

  • 31 Jul 2026 — HWE-Bench opens a leaderboard and tightens the run setup; GPT-5.6 Sol leads at 79.9%. Source
  • 16 Apr 2026 — Peking University releases HWE-Bench; GPT-5.4 with Codex CLI resolves 70.7% of hardware bugs. Source

Where it sits in the atlas

Work ladder: Direct evidence for Architecture and engineering. On the work ladder it counts as 79.9% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Design electrical or electronic systems or equipment; Evaluate designs, specifications, or other technical data; Test performance of equipment or systems.

Architecture and engineeringSoftware engineering (partial)

Hardware design and debuggingIssue resolution

Last checked 24 Sep 2026 against 6 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: HWE-Bench by Peking University (https://arxiv.org/abs/2604.14709).

Then, if you used this page:

Can Agents Work. "HWE-Bench: frontier results and sources." https://canagentswork.com/benchmarks/hwe-bench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-hwe-bench,
  title        = {{HWE-Bench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/hwe-bench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}