gold
SELECT DISTINCT language FROM foreign_data WHERE name = 'A Pedra Fellwar'
this statement states no ordering of its own
sha256:65ae88e2c50c13efc350eeb142b8f41e285c74a4e217a4ad445abf23e32ba0ca
R-SET NOT_EQUAL
card_games · mini_dev_sqlite from https://bird-bench.oss-cn-beijing.aliyuncs.com/minidev.zip (sha256 cc48ba16838204e4e214512030cb572eeb5f7bcdd999bae4b9b6ff12ec13b92f, downloaded 2026-09-07), member minidev/MINIDEV/mini_dev_sqlite.json
Which foreign language used by "A Pedra Fellwar"?
the hint the set supplies: "A Pedra Fellwar" refers to name = 'A Pedra Fellwar'
NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.
SELECT DISTINCT language FROM foreign_data WHERE name = 'A Pedra Fellwar'
this statement states no ordering of its own
sha256:65ae88e2c50c13efc350eeb142b8f41e285c74a4e217a4ad445abf23e32ba0ca
SELECT language
FROM foreign_data
WHERE name = 'A Pedra Fellwar';
this statement states no ordering of its own
sha256:8372ff671f7eb868037a8372f7811a420d59383b19396beb6fb8e2c3b91223b4
The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.
from counterexample.json, 0 rows, up to 25 shown per side
| side | languageTEXT |
|---|---|
| no rows |
from counterexample.json, 5 rows, up to 25 shown per side, 1 on this page
| side | times | languageTEXT |
|---|---|---|
| second | 5 | Portuguese (Brazil) |
set(second_rows) == set(gold_rows), REAL cells as Python float as sqlite3 returns them, so Python equality holds 1 == 1.0 == True as the benchmark's own scorer does
https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py
result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are SQLite's as sqlite3 returns them
ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py
The class states what makes these two results unequal under this rule, read off the two results and nothing else. It does not state which of the two statements is wrong.
from counterexample.json, 1 row
| languageTEXT |
|---|
| Portuguese (Brazil) |
from counterexample.json, 6 rows
| languageTEXT |
|---|
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
the statement states no top level ORDER BY
{
"heuristic": true,
"reason": "the statement states no top level ORDER BY"
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
the statement states no LIMIT
{
"heuristic": true,
"reason": "the statement states no LIMIT"
}
rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data
{
"heuristic": true,
"rule": "R-SET",
"baseline_result_hash": "sha256:65ae88e2c50c13efc350eeb142b8f41e285c74a4e217a4ad445abf23e32ba0ca",
"baseline_result": {
"columns": [
{
"name": "language",
"declared_type": "TEXT"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "str",
"value": "Portuguese (Brazil)"
}
]
],
"result_hash": "sha256:65ae88e2c50c13efc350eeb142b8f41e285c74a4e217a4ad445abf23e32ba0ca"
},
"planner_statistics": {},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"foreign_data"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {
"legalities": 427907
},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "equal",
"differs": false,
"result_hash": "sha256:65ae88e2c50c13efc350eeb142b8f41e285c74a4e217a4ad445abf23e32ba0ca",
"result": {
"columns": [
{
"name": "language",
"declared_type": "TEXT"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "str",
"value": "Portuguese (Brazil)"
}
]
],
"result_hash": "sha256:65ae88e2c50c13efc350eeb142b8f41e285c74a4e217a4ad445abf23e32ba0ca"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
}
}
SELECT DISTINCT language FROM foreign_data WHERE name = 'A Pedra Fellwar'
result_hash sha256:65ae88e2c50c13efc350eeb142b8f41e285c74a4e217a4ad445abf23e32ba0ca recomputed from this JSON: match
record_hash sha256:c19a312634fe2e98553b67a4d3f6f19e7456d84ff830c00473adfea53cfb180a recomputed from this JSON: match
from evidence-gold.json, 1 row
| languageTEXT |
|---|
| Portuguese (Brazil) |
SELECT language
FROM foreign_data
WHERE name = 'A Pedra Fellwar';
result_hash sha256:8372ff671f7eb868037a8372f7811a420d59383b19396beb6fb8e2c3b91223b4 recomputed from this JSON: match
record_hash sha256:d96178bb64b54a84e1f5514fafbabeb317a5a278b22122e6da151622e3d7d577 recomputed from this JSON: match
from evidence-second.json, 6 rows
| languageTEXT |
|---|
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
| Portuguese (Brazil) |
re-run this statement read-only against SQLite 3.53.4 | file=/private/tmp/attestql-runs/data/minidev/dev_databases/card_games/card_games.sqlite | size=261820416 under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET
re-run this statement read-only against SQLite 3.53.4 | file=/private/tmp/attestql-runs/data/minidev/dev_databases/card_games/card_games.sqlite | size=261820416 under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET