minidev-sqlite

9 groups 99 runs

The groups of this benchmark

One row per prediction file, and one run per database it names, because a connection here is one file. The numbers are the sums of that group's runs, which they can be because no question is in two of them. Every benchmark is indexed a level up.

group runs audited compared EQUAL NOT_EQUAL ERROR probes fired credited by BIRD and NOT_EQUAL
gpt-35-turbo 11 498 409 164 245 89 24 27
gpt-35-turbo-instruct 11 498 372 144 228 126 23 26
gpt-4 11 498 472 209 263 26 26 32
gpt-4-32k 11 498 460 205 255 38 25 32
gpt-4-turbo 11 498 422 190 232 24 29 76
meta-llama-3-70b-instruct-2 11 498 440 177 263 58 23 29
meta-llama-3-8b-instruct-2 11 498 311 187 207 104 16 21
mistralai-mixtral-8x7b-instru-4 11 498 248 250 157 91 14 16
phi-3-medium-128k-instruct-1 11 498 366 235 131 132 20 25