Benchmarks / HWE-Bench
HWE-Bench
Built by Peking University · Fan Cui, Hongyuan Hou, Zizhang Luo, Chenyun Yin, et al. · released 16 Apr 2026
417 hardware bug-repair tasks from Peking University, built from real bug-fix pull requests in six open-source chip projects (OpenTitan, Ibex, CVA6, Caliptra, XiangShan, Rocket Chip) written in Verilog, SystemVerilog, and Chisel. A coding agent gets a bug report inside a containerized copy of the repository and must fix the RTL so that the project's own simulation tests pass.
Frontier
79.9%
Resolved rate
GPT-5.6 Sol max (Codex CLI) · harness: Codex CLI
31 Jul 2026 · Source: Peking University (benchmark maintainers)
333 of 417 tasks resolved. The JSON stores the count; we computed the percent (79.856...) and rounded to one decimal as the site does. Precision 0.832; cost 1178.73 (unit not stated in the JSON; the README calls it an estimated token-equivalent cost). Added on 2026-07-31 with the leaderboard launch, per the README news. OpenTitan 184/245, Ibex 32/35, CVA6 33/35, Caliptra 15/16, XiangShan 42/54, Rocket Chip 27/32.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Resolved rate: the share of the 417 tasks where the agent's patch makes the failing simulation test pass (fail-to-pass) under the project's native flow (Verilator, or Synopsys VCS for OpenTitan). The agent receives a generated problem statement, not the original issue text, to avoid leaking the fix. The maintainers run each model with its usual agent CLI (Codex CLI, Claude Code, Kimi CLI, or OpenHands) with web search disabled, one attempt per task. The leaderboard also reports file-level precision of the patch and an estimated cost.
Resolved rate: Share of the 417 tasks resolved (the included simulation test passes after the agent's patch). Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli, repo
- Grading
- automated-tests
- Tasks
- 417
- Contamination
- Tasks come from public pull-request history, so the fixes exist online. The maintainers give the agent a rewritten problem statement and disable web search in the agent CLIs. The README is not consistent on network access: the 2026-07-31 news says generated tasks "now explicitly prohibit external access", but the setup section says they "allow public network access during agent installation and execution" and keep only the verifier offline.
- Reuse
- The task dataset on Hugging Face (henryen/hwe-bench) is tagged Apache 2.0. The code is Apache 2.0. Docker images for five projects are public; OpenTitan images are not distributed because that flow needs a Synopsys VCS license. (open-apache)
Limits to keep in mind
- OpenTitan supplies 245 of the 417 tasks and needs a commercial Synopsys VCS license. Without it, only the 172 Verilator-based tasks can be run. Source
- One run per row and no confidence intervals on the leaderboard. Rows mix the model and its agent CLI, so they are not clean model rankings. Source
- The 2026-07-31 update changed the run setup (new Harbor snapshot, no external access) and re-evaluated GPT-5.5, whose score moved from 79.4% to 76.5%. The README names no re-run for the other rows. Source
- The maintainers call their rows "reference scores", and the site describes no way to submit results. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Resolved rate | Date | Source |
|---|---|---|---|
| GPT-5.6 Sol max (Codex CLI) · frontier harness: Codex CLI | 79.9% | 31 Jul 2026 | Peking University · primary |
| Claude Opus 4.7 max (Claude Code) harness: Claude Code | 74.6% | 30 Apr 2026 | Peking University · primary |
| GPT-5.4 xhigh (Codex CLI) harness: Codex CLI | 70.7% | 16 Apr 2026 | Peking University · primary |
Timeline
Where it sits in the atlas
Work ladder: Direct evidence for Architecture and engineering. On the work ladder it counts as 79.9% of tasks completed. How the ladder works
Work it measures (O*NET work activities): Design electrical or electronic systems or equipment; Evaluate designs, specifications, or other technical data; Test performance of equipment or systems.
Go to the source
- Website pku-liang.github.io
- Paper arxiv.org
- Full leaderboard pku-liang.github.io
- Code github.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 6 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: HWE-Bench by Peking University (https://arxiv.org/abs/2604.14709).
Then, if you used this page:
Can Agents Work. "HWE-Bench: frontier results and sources." https://canagentswork.com/benchmarks/hwe-bench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-hwe-bench,
title = {{HWE-Bench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/hwe-bench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}