Public benchmarks
See how AI systems actually perform.
SPARKIT publishes independent tests of AI models and agents. Each release shows what we tested, how we scored it, where the systems failed, and what the results can and cannot tell you.
Benchmark releases
ReleaseWhat it testsResultsOpen
Release
ABI beta
What it tests
19 AI models on 115 difficult biology questions
Results
Overall scores, estimated cost, and response time
Release
HLE-Gold
What it tests
149 biology, medicine, and chemistry questions
Results
Accuracy, test method, supporting files, and known limits