Benchmarks / BIRD-Interact
BIRD-Interact
Built by HKU and Google Cloud · Nan Huo, Xiaohan Xu, Jinyang Li, Per Jacobsson, et al. · released 23 May 2025
Interactive text-to-SQL work on PostgreSQL databases with a simulated user. Requests are ambiguous, and the agent must ask the user, read database documentation and a knowledge base, write and debug SQL across create, read, update, and delete operations, and then handle a follow-up request.
Frontier
22.7%
Success rate, a-Interact, Full, both phases
MERIT + GPT-5.4 (GPT-4o user simulator) · agent: MERIT
22 May 2026 · Source: HKU (benchmark maintainers)
Rank 1 in the board data (a-Interact, Full, Stress mode, GPT-4o simulator), seen 2026-09-24. Self-submitted and marked verified:false; the institute logo file is named snowflake_logo. First-phase success 38.00%; Normalized Reward 33.45. In Free mode, an unverified AWS entry (Quick-BlendedIntelligence, Free-4) reached 20.50% on 2026-09-15.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Whether a database assistant can turn an underspecified business request into correct SQL through dialogue. The Full set has 600 tasks over 22 PostgreSQL databases (244 tables), each with executable test cases. Every task has two phases: resolve the ambiguities and deliver the first SQL, then handle a follow-up. In c-Interact the conversation follows a fixed protocol; in a-Interact the agent decides when to ask the user simulator, explore the database, or submit, and pays for each action from a fixed budget of "bird-coins". We show a-Interact on the Full set, Stress mode (the board's default), with success counted only when both phases pass.
Success rate, a-Interact, Full, both phases: Share of the 600 Full-set tasks where the agent's SQL passes the test cases for both the ambiguity-resolution phase and the follow-up phase in agentic mode ("Follow Ups" on the board). A completion rate. The board's default tab shows the first phase only ("Priority Questions"), and a Normalized Reward that also counts partial progress and budget use. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools, chat
- Grading
- automated-tests
- Tasks
- 600
- Contamination
- Gold SQL and test cases are not posted with the data; the maintainers send them by email. Databases come from LiveSQLBench, which the maintainers describe as contamination-free and evolving.
- Reuse
- CC BY-SA 4.0 (README). Task data is on Hugging Face; the gold SQL and test cases are sent by email on request to limit crawling. (open-other)
Limits to keep in mind
- Results depend on the user simulator (GPT-4o, Gemini-2.0-Flash, or Claude-Haiku-4-5). The board mixes simulators when "All" is selected, so entries are not strictly comparable across simulators. Source
- Some 2026 entries are self-submitted and marked unverified in the board data (verified: false); the maintainers verify a submission only after they run its code themselves. The current top entry, MERIT + GPT-5.4, is one of these. Source
- The board has three dataset versions (Mini 300 SQLite, Lite 300, Full 600), two interaction modes, Free and Stress budget modes, and several sub-metrics. Numbers quoted without that context (for example "38.00%") often refer to the first phase only. Source
- Docker database loads can fail silently and produce abnormally low scores; the maintainers ask users to check the logs before evaluation. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
- 26 Aug 2025 — BIRD team releases BIRD-Interact-Full; GPT-5 completes 17.00% of tasks. Source
Where it sits in the atlas
Work ladder: Direct evidence for Data and analytics. On the work ladder it counts as 22.7% of tasks completed. How the ladder works
Work it measures (O*NET work activities): Process digital or online data; Program computer systems or production equipment; Communicate with others about specifications or project details.
Go to the source
- Website bird-interact.github.io
- Paper arxiv.org
- Full leaderboard bird-interact.github.io
- Code github.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 4 primary sources. See an error? Tell us.
How to cite
Credit the original work first: BIRD-Interact by HKU and Google Cloud (https://arxiv.org/abs/2510.05318).
Then, if you used this page:
Can Agents Work. "BIRD-Interact: frontier results and sources." https://canagentswork.com/benchmarks/bird-interact/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-bird-interact,
title = {{BIRD-Interact: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/bird-interact/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}