Benchmarks / BIRD-Interact

BIRD-Interact

Built by HKU and Google Cloud · Nan Huo, Xiaohan Xu, Jinyang Li, Per Jacobsson, et al. · released 23 May 2025

Interactive text-to-SQL work on PostgreSQL databases with a simulated user. Requests are ambiguous, and the agent must ask the user, read database documentation and a knowledge base, write and debug SQL across create, read, update, and delete operations, and then handle a follow-up request.

Frontier

22.7%

Success rate, a-Interact, Full, both phases

MERIT + GPT-5.4 (GPT-4o user simulator) · agent: MERIT

22 May 2026 · Source: HKU (benchmark maintainers)

Rank 1 in the board data (a-Interact, Full, Stress mode, GPT-4o simulator), seen 2026-09-24. Self-submitted and marked verified:false; the institute logo file is named snowflake_logo. First-phase success 38.00%; Normalized Reward 33.45. In Free mode, an unverified AWS entry (Quick-BlendedIntelligence, Free-4) reached 20.50% on 2026-09-15.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at HKU

BIRD-Interact: Success rate, a-Interact, Full, both phases over time, 3 recorded results. 0%20%40%60%80%100%Oct 2025Jan 2026Apr 2026 GPT-5 (GPT-4o user simulator): 17% (22 Aug 2025) Claude-Opus-4.6 (Claude-Haiku-4-5 user simulator): 17.5% (17 Feb 2026) MERIT + GPT-5.4 (GPT-4o user simulator): 22.7% (22 May 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/bird-interact/"><img src="https://canagentswork.com/og/benchmarks-bird-interact.png" width="600" height="315" alt="BIRD-Interact: the best result is 22.7% (MERIT + GPT-5.4 (GPT-4o user simulator), 22 May 2026)." loading="lazy"></a>

Markdown:

[![BIRD-Interact: the best result is 22.7% (MERIT + GPT-5.4 (GPT-4o user simulator), 22 May 2026).](https://canagentswork.com/og/benchmarks-bird-interact.png)](https://canagentswork.com/benchmarks/bird-interact/)

What it measures

Whether a database assistant can turn an underspecified business request into correct SQL through dialogue. The Full set has 600 tasks over 22 PostgreSQL databases (244 tables), each with executable test cases. Every task has two phases: resolve the ambiguities and deliver the first SQL, then handle a follow-up. In c-Interact the conversation follows a fixed protocol; in a-Interact the agent decides when to ask the user simulator, explore the database, or submit, and pays for each action from a fixed budget of "bird-coins". We show a-Interact on the Full set, Stress mode (the board's default), with success counted only when both phases pass.

Success rate, a-Interact, Full, both phases: Share of the 600 Full-set tasks where the agent's SQL passes the test cases for both the ambiguity-resolution phase and the follow-up phase in agentic mode ("Follow Ups" on the board). A completion rate. The board's default tab shows the first phase only ("Priority Questions"), and a Normalized Reward that also counts partial progress and budget use. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, chat
Grading
automated-tests
Tasks
600
Contamination
Gold SQL and test cases are not posted with the data; the maintainers send them by email. Databases come from LiveSQLBench, which the maintainers describe as contamination-free and evolving.
Reuse
CC BY-SA 4.0 (README). Task data is on Hugging Face; the gold SQL and test cases are sent by email on request to limit crawling. (open-other)

Limits to keep in mind

  • Results depend on the user simulator (GPT-4o, Gemini-2.0-Flash, or Claude-Haiku-4-5). The board mixes simulators when "All" is selected, so entries are not strictly comparable across simulators. Source
  • Some 2026 entries are self-submitted and marked unverified in the board data (verified: false); the maintainers verify a submission only after they run its code themselves. The current top entry, MERIT + GPT-5.4, is one of these. Source
  • The board has three dataset versions (Mini 300 SQLite, Lite 300, Full 600), two interaction modes, Free and Stress budget modes, and several sub-metrics. Numbers quoted without that context (for example "38.00%") often refer to the first phase only. Source
  • Docker database loads can fail silently and produce abnormally low scores; the maintainers ask users to check the logs before evaluation. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemSuccess rate, a-Interact, Full, both phasesDateSource
MERIT + GPT-5.4 (GPT-4o user simulator) · frontier22.7%22 May 2026HKU · primary
Claude-Opus-4.6 (Claude-Haiku-4-5 user simulator)17.5%17 Feb 2026HKU · primary
GPT-5 (GPT-4o user simulator)17%22 Aug 2025HKU · primary

Timeline

  • 26 Aug 2025 — BIRD team releases BIRD-Interact-Full; GPT-5 completes 17.00% of tasks. Source

Where it sits in the atlas

Work ladder: Direct evidence for Data and analytics. On the work ladder it counts as 22.7% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Process digital or online data; Program computer systems or production equipment; Communicate with others about specifications or project details.

Data and analyticsSoftware engineering (partial)

Data engineering and SQLAsking for clarification

Last checked 24 Sep 2026 against 4 primary sources. See an error? Tell us.

How to cite

Credit the original work first: BIRD-Interact by HKU and Google Cloud (https://arxiv.org/abs/2510.05318).

Then, if you used this page:

Can Agents Work. "BIRD-Interact: frontier results and sources." https://canagentswork.com/benchmarks/bird-interact/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-bird-interact,
  title        = {{BIRD-Interact: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/bird-interact/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}