gold
SELECT DISTINCT T.element FROM atom AS T WHERE T.molecule_id = 'TR004'
this statement states no ordering of its own
sha256:bf223b38e0296761b318c97ba13cec1cc74e9b3e2579671077dc3364ca24d5ee
R-SET NOT_EQUAL
toxicology · mini_dev_postgresql from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_postgresql.json
List all the elements of the toxicology of the molecule "TR004".
the hint the set supplies: TR004 is the molecule id;
NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.
read by hand, 2026-09-04: A a wrong answer the benchmark credited: another table or projection; a scalar repeated per row of an unrelated table; a list at least twice the answer's length where the question asks for the things, not the rows
gold DISTINCT over 6 elements; the prediction lists every atom, 24 rows
a maintainer's reading of this question, out of classification.json, copied from plans/reports/prediction-mode-260904-real-predictions/classification.json. It is not a verdict and nothing above it was computed from it.
SELECT DISTINCT T.element FROM atom AS T WHERE T.molecule_id = 'TR004'
this statement states no ordering of its own
sha256:bf223b38e0296761b318c97ba13cec1cc74e9b3e2579671077dc3364ca24d5ee
SELECT element
FROM atom
WHERE molecule_id = 'TR004';
this statement states no ordering of its own
sha256:aa823261b57a7d4c10e1280cf86f26719f159697eac1b1f49f38e256d30c1cf9
The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.
from counterexample.json, 0 rows, up to 25 shown per side
| side | elementtext |
|---|---|
| no rows |
from counterexample.json, 18 rows, up to 25 shown per side, 4 on this page
| side | times | elementtext |
|---|---|---|
| second | 11 | h |
| second | 4 | c |
| second | 2 | o |
| second | 1 | s |
set(second_rows) == set(gold_rows), float4/float8 cells as Python float as psycopg2 returns them, numeric as Decimal
https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py
result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are PostgreSQL's as psycopg2 returns them
ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py
The class states what makes these two results unequal under this rule, read off the two results and nothing else. It does not state which of the two statements is wrong.
from counterexample.json, 6 rows
| elementtext |
|---|
| c |
| h |
| n |
| o |
| p |
| s |
from counterexample.json, 24 rows
| elementtext |
|---|
| s |
| n |
| o |
| c |
| h |
| h |
| h |
| h |
| h |
| h |
| h |
| p |
| h |
| h |
| h |
| h |
| h |
| s |
| o |
| o |
| c |
| c |
| c |
| c |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
the statement states no top level ORDER BY
{
"heuristic": true,
"reason": "the statement states no top level ORDER BY"
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
the statement states no LIMIT
{
"heuristic": true,
"reason": "the statement states no LIMIT"
}
rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data
{
"heuristic": true,
"rule": "R-SET",
"baseline_result_hash": "sha256:bf223b38e0296761b318c97ba13cec1cc74e9b3e2579671077dc3364ca24d5ee",
"baseline_result": {
"columns": [
{
"name": "element",
"declared_type": "text"
}
],
"row_count": 6,
"truncated": false,
"rows_shown": 6,
"rows": [
[
{
"type": "str",
"value": "c"
}
],
[
{
"type": "str",
"value": "h"
}
],
[
{
"type": "str",
"value": "n"
}
],
[
{
"type": "str",
"value": "o"
}
],
[
{
"type": "str",
"value": "p"
}
],
[
{
"type": "str",
"value": "s"
}
]
],
"result_hash": "sha256:bf223b38e0296761b318c97ba13cec1cc74e9b3e2579671077dc3364ca24d5ee"
},
"planner_statistics": {
"atom": {
"last_analyze": null,
"last_autoanalyze": "2026-09-08 05:19:10.04421+00",
"n_mod_since_analyze": 0
}
},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"atom"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {
"laptimes": 400524,
"legalities": 427907,
"posthistory": 303155,
"trans": 1056320,
"yearmonth": 383282
},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "equal",
"differs": false,
"result_hash": "sha256:bf223b38e0296761b318c97ba13cec1cc74e9b3e2579671077dc3364ca24d5ee",
"result": {
"columns": [
{
"name": "element",
"declared_type": "text"
}
],
"row_count": 6,
"truncated": false,
"rows_shown": 6,
"rows": [
[
{
"type": "str",
"value": "c"
}
],
[
{
"type": "str",
"value": "h"
}
],
[
{
"type": "str",
"value": "n"
}
],
[
{
"type": "str",
"value": "o"
}
],
[
{
"type": "str",
"value": "p"
}
],
[
{
"type": "str",
"value": "s"
}
]
],
"result_hash": "sha256:bf223b38e0296761b318c97ba13cec1cc74e9b3e2579671077dc3364ca24d5ee"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
}
}
SELECT DISTINCT T.element FROM atom AS T WHERE T.molecule_id = 'TR004'
result_hash sha256:bf223b38e0296761b318c97ba13cec1cc74e9b3e2579671077dc3364ca24d5ee recomputed from this JSON: match
record_hash sha256:b4387bdd02108fba4952fcb19baf7c3fe47b53e8a70d92444fee98be10b59b95 recomputed from this JSON: match
from evidence-gold.json, 6 rows
| elementtext |
|---|
| c |
| h |
| n |
| o |
| p |
| s |
SELECT element
FROM atom
WHERE molecule_id = 'TR004';
result_hash sha256:aa823261b57a7d4c10e1280cf86f26719f159697eac1b1f49f38e256d30c1cf9 recomputed from this JSON: match
record_hash sha256:c14a7624cd37a68e8db58312dd5164c21663c3b953105bda0d7b5e1790dd86dd recomputed from this JSON: match
from evidence-second.json, 24 rows
| elementtext |
|---|
| s |
| n |
| o |
| c |
| h |
| h |
| h |
| h |
| h |
| h |
| h |
| p |
| h |
| h |
| h |
| h |
| h |
| s |
| o |
| o |
| c |
| c |
| c |
| c |
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET