Measured August 2026
HLE-Gold
SPARKIT answered 64.4% of the 149 evaluated HLE-Gold questions correctly, compared with 53.0% for direct Claude Opus 5 and 39.0% for direct GPT-5.6-Sol.
149 questions · Accuracy · Updated August 11, 2026
Results
| System | Test setup | Accuracy |
|---|---|---|
| SPARKITSPARKIT research-agent run | SPARKIT research-agent run | 64.4% |
| Claude Opus 5Direct model call | Direct model call | 53.0% |
| GPT-5.6-SolDirect model call | Direct model call | 39.0% |
How we ran the test
- 01The evaluation uses 149 biology, medicine, and chemistry questions from the gold subset of Humanity's Last Exam.
- 02Each system is scored against the benchmark answer key, and the reported value is the percentage of evaluated questions answered correctly.
- 03The direct-model baselines are intentionally shown without the SPARKIT search, reading, and analysis loop; they measure the value of the research workflow rather than a model-family comparison.
Sources and supporting files
Humanity's Last Exam paper ↗The primary publication describing the public benchmark.Representative HLE-Gold run artifacts (JSON) →Three difficult example questions with run metadata and outputs. These examples are not the complete benchmark dataset.End-to-end HLE-Gold walkthrough →A worked example showing the question, research trace, report, and grading outcome.
Limitations
- 01HLE-Gold is a difficult scientific subset, not a comprehensive measure of every research domain or production workflow.
- 02Model and agent behavior can change as providers update their systems, so these results are a dated snapshot rather than a permanent ranking.
- 03Representative run artifacts are published below, but the complete run-level benchmark output is not yet public.
- 04Benchmark accuracy is not evidence of clinical, legal, or regulatory reliability; cited sources and conclusions still require expert verification.