SPARKIT
← All benchmarks

Measured August 2026

HLE-Gold

SPARKIT answered 64.4% of the 149 evaluated HLE-Gold questions correctly, compared with 53.0% for direct Claude Opus 5 and 39.0% for direct GPT-5.6-Sol.

149 questions · Accuracy · Updated August 11, 2026

Results

SystemAccuracy
SPARKITSPARKIT research-agent run64.4%
Claude Opus 5Direct model call53.0%
GPT-5.6-SolDirect model call39.0%

How we ran the test

  1. 01The evaluation uses 149 biology, medicine, and chemistry questions from the gold subset of Humanity's Last Exam.
  2. 02Each system is scored against the benchmark answer key, and the reported value is the percentage of evaluated questions answered correctly.
  3. 03The direct-model baselines are intentionally shown without the SPARKIT search, reading, and analysis loop; they measure the value of the research workflow rather than a model-family comparison.

Sources and supporting files

Limitations

  1. 01HLE-Gold is a difficult scientific subset, not a comprehensive measure of every research domain or production workflow.
  2. 02Model and agent behavior can change as providers update their systems, so these results are a dated snapshot rather than a permanent ranking.
  3. 03Representative run artifacts are published below, but the complete run-level benchmark output is not yet public.
  4. 04Benchmark accuracy is not evidence of clinical, legal, or regulatory reliability; cited sources and conclusions still require expert verification.