ABI beta · early preview
Artificial Biological Intelligence Index
ABI is a public test of how well AI models answer difficult biology questions. This early release compares nineteen models on three test sets: LABBench2, HLE-Gold, and CGBench. It also shows how each model's score relates to cost and response time.
- models tested
- 19
- questions
- 115
- test sets
- 3
Overall scores · v0.2
How the models scored overall.
Each bar shows one model's combined ABI score across all 115 questions. A taller bar means a higher score on this test.
Cost and speed · v0.2.1
Better answers can cost more and take longer.
The left chart compares score with estimated cost per question. The right chart compares score with response time. Higher is better; farther left is cheaper or faster. The dashed line marks models that deliver the highest score at a given cost or speed.
This chart does not pick the best model for every job. Your choice still depends on the work, the mistakes you can tolerate, the provider's rules, how you plan to run the model, and how cost and response time were measured.