All benchmarks

ABI beta · early preview

Artificial Biological Intelligence Index

ABI is a public test of how well AI models answer difficult biology questions. This early release compares nineteen models on three test sets: LABBench2, HLE-Gold, and CGBench. It also shows how each model's score relates to cost and response time.

models tested
19
questions
115
test sets
3

Overall scores · v0.2

How the models scored overall.

Each bar shows one model's combined ABI score across all 115 questions. A taller bar means a higher score on this test.

Bar chart comparing nineteen AI models on 115 biology questions from LABBench2, HLE-Gold, and CGBench.
These early scores were graded automatically and may change. Each of the 115 questions counts the same. Read how the test was run and what it cannot prove before drawing conclusions.

Cost and speed · v0.2.1

Better answers can cost more and take longer.

The left chart compares score with estimated cost per question. The right chart compares score with response time. Higher is better; farther left is cheaper or faster. The dashed line marks models that deliver the highest score at a given cost or speed.

Two charts comparing ABI scores with cost and response time for nineteen AI models. Dashed lines show the strongest score and cost or speed combinations.
Four Muse answers had no price data. We only know its minimum possible cost, so Muse is left off the dashed cost line.Download SVG

This chart does not pick the best model for every job. Your choice still depends on the work, the mistakes you can tolerate, the provider's rules, how you plan to run the model, and how cost and response time were measured.