gold
SELECT SUM(cost) FROM expense WHERE expense_description = 'Pizza'
this statement states no ordering of its own
sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84
R-SET GOLD-ONLY float-aggregate-order
student_club · mini_dev_pg-00000-of-00001 from https://huggingface.co/datasets/birdsql/bird_mini_dev/resolve/f65faf4ae3b638c1fa6df1d3370c8d92c8366301/data/mini_dev_pg-00000-of-00001.json (commit f65faf4a, downloaded 2026-09-08)
What is the total cost of the pizzas for all the events?
the hint the set supplies: total cost of the pizzas refers to SUM(cost) where expense_description = 'Pizza'
This question was audited without a prediction beside it, so there is nothing to compare the gold with. The probes below read the gold alone.
SELECT SUM(cost) FROM expense WHERE expense_description = 'Pizza'
this statement states no ordering of its own
sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84
from evidence-gold.json, 1 row
| sumfloat4 |
|---|
| 600.11 |
A smell is a mechanical reason to read this gold statement again. It is a heuristic: it does not state that the statement is wrong, and a maintainer decides.
this statement orders by a text column holding only numbers, and ordering it as a number gives a different answer, so the gold may be sorting 9.5 above 10
the statement states no top level ORDER BY
{
"heuristic": true,
"reason": "the statement states no top level ORDER BY"
}
this statement cuts its result at a LIMIT that does not decide which rows come back, so a different but equally correct statement can return other rows and score zero
the statement states no LIMIT
{
"heuristic": true,
"reason": "the statement states no LIMIT"
}
this statement aggregates floating point numbers, so its last digits depend on the order the rows were summed in; the values agree to six significant digits
from smells.json, 1 row
| 600.11005 |
{
"heuristic": true,
"rule": "R-SET",
"baseline_result_hash": "sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84",
"baseline_result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "600.11"
}
]
],
"result_hash": "sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84"
},
"planner_statistics": {
"expense": {
"last_analyze": null,
"last_autoanalyze": null,
"n_mod_since_analyze": 32
}
},
"shuffle": {
"seed": "1",
"row_limit": 300000,
"tables": [
"expense"
],
"tables_not_shuffled": [],
"tables_skipped_for_size": {
"laptimes": 400524,
"legalities": 427907,
"posthistory": 303155,
"trans": 1056320,
"yearmonth": 383282
},
"tables_not_reached_by_a_copy": {}
},
"shuffled_copies": {
"run": true,
"verdict": "not_equal",
"differs": true,
"result_hash": "sha256:28a33fd0933dc8c2b7eb34dc1b147cee178aa7665999cce122d919d75a4750d2",
"result": {
"columns": [
{
"name": "sum",
"declared_type": "float4"
}
],
"row_count": 1,
"truncated": false,
"rows_shown": 1,
"rows": [
[
{
"type": "dec",
"value": "600.11005"
}
]
],
"result_hash": "sha256:28a33fd0933dc8c2b7eb34dc1b147cee178aa7665999cce122d919d75a4750d2"
}
},
"plan_variant": {
"run": false,
"reason": "the plan variant was not asked for"
},
"float_cells": [
{
"row": 0,
"column": "sum",
"declared_type": "float4",
"baseline": "600.11",
"rerun": "600.11005"
}
],
"significant_digits": 6
}
SELECT SUM(cost) FROM expense WHERE expense_description = 'Pizza'
result_hash sha256:0755eee7cb5f3b509bba8df2e532e80cca0094942fe7905c819842bbede6fd84 recomputed from this JSON: match
record_hash sha256:5e4dbc4bf8338deb8cd6f44aaab5f10eb02e08aabe15828910cb3a964816ebca recomputed from this JSON: match
from evidence-gold.json, 1 row
| sumfloat4 |
|---|
| 600.11 |
re-run this statement read-only against PostgreSQL 16.15 (Debian 16.15-1.pgdg13+2) on aarch64-unknown-linux-gnu, compiled by gcc (Debian 14.2.0-19) 14.2.0, 64-bit | server=172.17.0.2/32:5432 | database=bird under the session settings and over the data this record's fixture digest names, and compare the two results under R-SET