SPARKIT

Public evaluations

Research agent benchmarks

Headline results are useful only when the evaluated questions, exact comparators, supporting evidence, and limitations remain attached. This page is the canonical index for SPARKIT's published benchmark claims.

Measured June 2026 · Results updated July 15, 2026

Accuracy

HLE-Gold

149 questions · biology / medicine + chemistry

SPARKIT answered 54.4% of the 149 evaluated HLE-Gold questions correctly, compared with 39.0% for direct GPT-5.6-Sol and 34.9% for direct Claude Opus 4.8.

Methods, evidence, and limitations →
HLE-Gold benchmark results
SystemAccuracy
SPARKITSPARKIT research-agent run54.4%
GPT-5.6-SolDirect model call39.0%
Claude Opus 4.8Direct model call34.9%

Accuracy

GAIA

127 questions

SPARKIT answered 73.2% of the 127 evaluated GAIA questions correctly, compared with 58.2% for Exa and 57.0% for Brave.

Methods, evidence, and limitations →
GAIA benchmark results
SystemAccuracy
SPARKITSPARKIT research-agent run73.2%
ExaSearch API result58.2%
BraveSearch API result57.0%