q207

R-SET NOT_EQUAL

questions

The question

What elements are in a double type bond?

the hint the set supplies: double type bond refers to bond_type = '=';

NOT_EQUAL states that these two statements disagree on this data under this rule. It does not state which of them is wrong.

The statements

gold

SELECT DISTINCT T1.element FROM atom AS T1 INNER JOIN bond AS T2 ON T1.molecule_id = T2.molecule_id INNER JOIN connected AS T3 ON T1.atom_id = T3.atom_id WHERE T2.bond_type = '='

this statement states no ordering of its own

sha256:029fb5dbc54b4b7234684a95a9d6e26ec41b884c6e7ce239fabb69a0ae9d6764

second

SELECT DISTINCT T1.element FROM atom AS T1 JOIN connected AS T2 ON T1.atom_id = T2.atom_id JOIN bond AS T3 ON T2.bond_id = T3.bond_id WHERE T3.bond_type = '='

this statement states no ordering of its own

sha256:9d2e4c969dafd20791a3829f3bbd056403b619bcedd6a6105e58228c98fa2bee

The marked tokens are where the two texts differ. Two statements that differ everywhere can return the same rows, and two that differ in one token can return other rows; the verdict above is read off the results.

The rows they differ in

in gold, not in the prediction, 1 row

from counterexample.json, 1 row, up to 25 shown per side

side elementTEXT
gold n

in the prediction, not in gold, 0 rows

from counterexample.json, 0 rows, up to 25 shown per side

side elementTEXT
no rows

What the benchmark would have said

BIRD's own check: 0

set(second_rows) == set(gold_rows), REAL cells as Python float as sqlite3 returns them, so Python equality holds 1 == 1.0 == True as the benchmark's own scorer does

https://github.com/bird-bench/mini_dev/blob/main/evaluation/evaluation_ex.py

gold_rows
3
second_rows
2
gold_distinct_rows
3
second_distinct_rows
2

the test-suite check: 0

result_eq: equal row counts and equal column counts, each row unordered as a quick rejection, then the two equal as a list when the gold text holds ORDER BY and as a multiset otherwise, under some permutation of the columns; DISTINCT is not stripped and re-executed, and the cells are SQLite's as sqlite3 returns them

ruiqi-zhong/test-suite-sql-eval, exec_eval.py, result_eq, at commit 48cb78ec: https://github.com/ruiqi-zhong/test-suite-sql-eval/blob/48cb78ecf7f610620206283846c76751b18a1326/exec_eval.py

order_matters
false
gold_rows
3
second_rows
2
gold_columns
1
second_columns
1

the class: truncation

The class states what makes these two results unequal under this rule, read off the two results and nothing else. It does not state which of the two statements is wrong.

one result and the shorter one that is its first rows, from evidence-gold.json and evidence-second.json3 gold2 prediction
The gold returned 3 rows and the prediction 2, which are the first 2 rows of the other one in the same order.
gold_types
["TEXT"]
second_types
["TEXT"]
multiset_equal
false
set_equal
false
order_equal
false
shorter_result_is_a_prefix
true

The results

gold, 3 rows

from counterexample.json, 3 rows

elementTEXT
c
o
n
second, 2 rows

from counterexample.json, 2 rows

elementTEXT
c
o

The probes

A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.

ordering-over-numeric-text not applicable

this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10

the statement states no top level ORDER BY

what it measured
{
  "heuristic": true,
  "reason": "the statement states no top level ORDER BY"
}

arbitrary-cut not applicable

this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero

the statement states no LIMIT

what it measured
{
  "heuristic": true,
  "reason": "the statement states no LIMIT"
}

not-a-function-of-the-data quiet

rerun over the same rows in another physical order this statement gives another answer, so its result depends on how the rows are stored and not only on the data

what it measured
{
  "heuristic": true,
  "rule": "R-SET",
  "baseline_result_hash": "sha256:029fb5dbc54b4b7234684a95a9d6e26ec41b884c6e7ce239fabb69a0ae9d6764",
  "baseline_result": {
    "columns": [
      {
        "name": "element",
        "declared_type": "TEXT"
      }
    ],
    "row_count": 3,
    "truncated": false,
    "rows_shown": 3,
    "rows": [
      [
        {
          "type": "str",
          "value": "c"
        }
      ],
      [
        {
          "type": "str",
          "value": "o"
        }
      ],
      [
        {
          "type": "str",
          "value": "n"
        }
      ]
    ],
    "result_hash": "sha256:029fb5dbc54b4b7234684a95a9d6e26ec41b884c6e7ce239fabb69a0ae9d6764"
  },
  "planner_statistics": {},
  "shuffle": {
    "seed": "1",
    "row_limit": 300000,
    "tables": [
      "atom",
      "bond",
      "connected"
    ],
    "tables_not_shuffled": [],
    "tables_skipped_for_size": {},
    "tables_not_reached_by_a_copy": {
      "main.y": "the statement names this table's schema, and a name that states its schema is read from that schema whatever TEMP holds, so the rerun reads this table and not a copy of it"
    }
  },
  "shuffled_copies": {
    "run": true,
    "verdict": "equal",
    "differs": false,
    "result_hash": "sha256:029fb5dbc54b4b7234684a95a9d6e26ec41b884c6e7ce239fabb69a0ae9d6764",
    "result": {
      "columns": [
        {
          "name": "element",
          "declared_type": "TEXT"
        }
      ],
      "row_count": 3,
      "truncated": false,
      "rows_shown": 3,
      "rows": [
        [
          {
            "type": "str",
            "value": "c"
          }
        ],
        [
          {
            "type": "str",
            "value": "o"
          }
        ],
        [
          {
            "type": "str",
            "value": "n"
          }
        ]
      ],
      "result_hash": "sha256:029fb5dbc54b4b7234684a95a9d6e26ec41b884c6e7ce239fabb69a0ae9d6764"
    }
  },
  "plan_variant": {
    "run": false,
    "reason": "the plan variant was not asked for"
  }
}

The evidence records

gold: evidence-gold.json

SELECT DISTINCT T1.element FROM atom AS T1 INNER JOIN bond AS T2 ON T1.molecule_id = T2.molecule_id INNER JOIN connected AS T3 ON T1.atom_id = T3.atom_id WHERE T2.bond_type = '='
statement read from
demo/questions.json
digest
sha256:7ad58addc070b425cabfc9d67bbd9c95840253160cf3b73d672e9aeb4897f5ff
origin
the run was told none

result_hash sha256:029fb5dbc54b4b7234684a95a9d6e26ec41b884c6e7ce239fabb69a0ae9d6764 recomputed from this JSON: match

record_hash sha256:8fa0613e80eadb7a86bab18324390df7f5fe73f783a8872be232188e44b5b1d4 recomputed from this JSON: match

the result this record holds, 3 rows

from evidence-gold.json, 3 rows

elementTEXT
c
o
n
what ran, and where
run
audit-5ac1b5ac-ceba-418a-841e-577381666b11
executed at
2026-09-08T04:11:01.299359+00:00
data as of
2026-09-08T04:11:01.285823+00:00
backend at checkout
SQLite 3.53.4 | file=/private/tmp/attestql-site-sandbox/demo/fixture.sqlite | size=69632
backend that answered
SQLite 3.53.4 | file=/private/tmp/attestql-site-sandbox/demo/fixture.sqlite | size=69632
database role
file
replay rule
R-SET
question set version
sha256:7ad58addc070b425cabfc9d67bbd9c95840253160cf3b73d672e9aeb4897f5ff
validator
audit:sqlglot-sqlite-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
30000 ms
rows
3 rows
the session it ran under
engine
sqlite
time_zone
not stated by this engine
date_style
not stated by this engine
interval_style
not stated by this engine
extra_float_digits
not stated by this engine
database_collation
not stated by this engine
work_mem
not stated by this engine
hash_mem_multiplier
not stated by this engine

recorded beside them

sqlite_version
3.53.4
encoding
UTF-8
reverse_unordered_selects
0
query_only
1
journal_mode
delete
data_version
1
compile_options
ATOMIC_INTRINSICS=1,COMPILER=clang-21.0.0,DEFAULT_AUTOVACUUM,DEFAULT_CACHE_SIZE=-2000,DEFAULT_FILE_FORMAT=4,DEFAULT_JOURNAL_SIZE_LIMIT=-1,DEFAULT_MMAP_SIZE=0,DEFAULT_PAGE_SIZE=4096,DEFAULT_PCACHE_INITSZ=20,DEFAULT_RECURSIVE_TRIGGERS,DEFAULT_SECTOR_SIZE=4096,DEFAULT_SYNCHRONOUS=2,DEFAULT_WAL_AUTOCHECKPOINT=1000,DEFAULT_WAL_SYNCHRONOUS=2,DEFAULT_WORKER_THREADS=0,DIRECT_OVERFLOW_READ,ENABLE_API_ARMOR,ENABLE_COLUMN_METADATA,ENABLE_DBSTAT_VTAB,ENABLE_FTS3,ENABLE_FTS3_PARENTHESIS,ENABLE_FTS5,ENABLE_GEOPOLY,ENABLE_MATH_FUNCTIONS,ENABLE_MEMORY_MANAGEMENT,ENABLE_PERCENTILE,ENABLE_PREUPDATE_HOOK,ENABLE_RTREE,ENABLE_SESSION,ENABLE_STAT4,ENABLE_UNLOCK_NOTIFY,MALLOC_SOFT_LIMIT=1024,MAX_ATTACHED=10,MAX_COLUMN=2000,MAX_COMPOUND_SELECT=500,MAX_DEFAULT_PAGE_SIZE=8192,MAX_EXPR_DEPTH=1000,MAX_FUNCTION_ARG=1000,MAX_LENGTH=1000000000,MAX_LIKE_PATTERN_LENGTH=50000,MAX_MMAP_SIZE=0x7fff0000,MAX_PAGE_COUNT=0xfffffffe,MAX_PAGE_SIZE=65536,MAX_SQL_LENGTH=1000000000,MAX_TRIGGER_DEPTH=1000,MAX_VARIABLE_NUMBER=250000,MAX_VDBE_OP=250000000,MAX_WORKER_THREADS=8,MUTEX_PTHREADS,SYSTEM_MALLOC,TEMP_STORE=1,THREADSAFE=1,USE_URI
collation_list
RTRIM,NOCASE,BINARY
case_sensitive_like
0
the rendering and the data
version
attestql/audit/2
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:c0d10ad3b9e4119035fdf3bf953e3693755ec5bfabd479c79d8a31d73407c19e
source file sha256
rows in main.atom
5
rows in main.bond
3
rows in main.connected
6

second: evidence-second.json

SELECT DISTINCT T1.element FROM atom AS T1 JOIN connected AS T2 ON T1.atom_id = T2.atom_id JOIN bond AS T3 ON T2.bond_id = T3.bond_id WHERE T3.bond_type = '='
statement read from
demo/predictions.json
digest
sha256:497ab56b10daa81359a51f3be419fe16e5b796b0cc33cb1aff6c24c9aa853780
origin
the run was told none

result_hash sha256:9d2e4c969dafd20791a3829f3bbd056403b619bcedd6a6105e58228c98fa2bee recomputed from this JSON: match

record_hash sha256:ab35b97982375492955a541eec9032c857d8e95c73b30ef52a7495337556bc5e recomputed from this JSON: match

the result this record holds, 2 rows

from evidence-second.json, 2 rows

elementTEXT
c
o
what ran, and where
run
audit-5ac1b5ac-ceba-418a-841e-577381666b11
executed at
2026-09-08T04:11:01.299522+00:00
data as of
2026-09-08T04:11:01.285823+00:00
backend at checkout
SQLite 3.53.4 | file=/private/tmp/attestql-site-sandbox/demo/fixture.sqlite | size=69632
backend that answered
SQLite 3.53.4 | file=/private/tmp/attestql-site-sandbox/demo/fixture.sqlite | size=69632
database role
file
replay rule
R-SET
question set version
sha256:7ad58addc070b425cabfc9d67bbd9c95840253160cf3b73d672e9aeb4897f5ff
validator
audit:sqlglot-sqlite-parse
checks run
parses_as_exactly_one_statement, the_one_statement_is_a_select, no_placeholder_without_a_bound_parameter
statement timeout
30000 ms
rows
2 rows
the session it ran under
engine
sqlite
time_zone
not stated by this engine
date_style
not stated by this engine
interval_style
not stated by this engine
extra_float_digits
not stated by this engine
database_collation
not stated by this engine
work_mem
not stated by this engine
hash_mem_multiplier
not stated by this engine

recorded beside them

sqlite_version
3.53.4
encoding
UTF-8
reverse_unordered_selects
0
query_only
1
journal_mode
delete
data_version
1
compile_options
ATOMIC_INTRINSICS=1,COMPILER=clang-21.0.0,DEFAULT_AUTOVACUUM,DEFAULT_CACHE_SIZE=-2000,DEFAULT_FILE_FORMAT=4,DEFAULT_JOURNAL_SIZE_LIMIT=-1,DEFAULT_MMAP_SIZE=0,DEFAULT_PAGE_SIZE=4096,DEFAULT_PCACHE_INITSZ=20,DEFAULT_RECURSIVE_TRIGGERS,DEFAULT_SECTOR_SIZE=4096,DEFAULT_SYNCHRONOUS=2,DEFAULT_WAL_AUTOCHECKPOINT=1000,DEFAULT_WAL_SYNCHRONOUS=2,DEFAULT_WORKER_THREADS=0,DIRECT_OVERFLOW_READ,ENABLE_API_ARMOR,ENABLE_COLUMN_METADATA,ENABLE_DBSTAT_VTAB,ENABLE_FTS3,ENABLE_FTS3_PARENTHESIS,ENABLE_FTS5,ENABLE_GEOPOLY,ENABLE_MATH_FUNCTIONS,ENABLE_MEMORY_MANAGEMENT,ENABLE_PERCENTILE,ENABLE_PREUPDATE_HOOK,ENABLE_RTREE,ENABLE_SESSION,ENABLE_STAT4,ENABLE_UNLOCK_NOTIFY,MALLOC_SOFT_LIMIT=1024,MAX_ATTACHED=10,MAX_COLUMN=2000,MAX_COMPOUND_SELECT=500,MAX_DEFAULT_PAGE_SIZE=8192,MAX_EXPR_DEPTH=1000,MAX_FUNCTION_ARG=1000,MAX_LENGTH=1000000000,MAX_LIKE_PATTERN_LENGTH=50000,MAX_MMAP_SIZE=0x7fff0000,MAX_PAGE_COUNT=0xfffffffe,MAX_PAGE_SIZE=65536,MAX_SQL_LENGTH=1000000000,MAX_TRIGGER_DEPTH=1000,MAX_VARIABLE_NUMBER=250000,MAX_VDBE_OP=250000000,MAX_WORKER_THREADS=8,MUTEX_PTHREADS,SYSTEM_MALLOC,TEMP_STORE=1,THREADSAFE=1,USE_URI
collation_list
RTRIM,NOCASE,BINARY
case_sensitive_like
0
the rendering and the data
version
attestql/audit/2
numeric_scale
6
timestamp_format
%Y-%m-%dT%H:%M:%S.%fZ
timezone
UTC
null_rendering
NULL
encoding
utf-8
schema digest
sha256:c0d10ad3b9e4119035fdf3bf953e3693755ec5bfabd479c79d8a31d73407c19e
source file sha256
rows in main.atom
5
rows in main.bond
3
rows in main.connected
6

Running these again

gold

re-run this statement read-only against SQLite 3.53.4 | file=/private/tmp/attestql-site-sandbox/demo/fixture.sqlite | size=69632 under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET

second

re-run this statement read-only against SQLite 3.53.4 | file=/private/tmp/attestql-site-sandbox/demo/fixture.sqlite | size=69632 under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET

This question's run

run
audit-5ac1b5ac-ceba-418a-841e-577381666b11
server
SQLite 3.53.4 | file=/private/tmp/attestql-site-sandbox/demo/fixture.sqlite | size=69632
question set
questions
replay rule
R-SET

the run this question belongs to

The JSON this page was rendered from