# SPARKIT > SPARKIT independently evaluates AI models, agents, and products, publishes public benchmarks, and helps teams integrate AI systems into real work. ## Services - **Independent AI evaluations:** SPARKIT builds tests around a customer's real tasks, runs the tests independently, and reports what works, what fails, and why. - **AI and agent integration:** SPARKIT connects selected AI systems to a team's existing software, data, tools, permissions, and review process. - **Research agent:** SPARKIT also offers an API and dashboard agent that searches the web and scientific literature, reads evidence, runs calculations or code, and returns cited reports. ## Research agent capabilities - Search web sources and scientific literature - Read papers, PDFs, and web pages - Compare evidence and identify caveats - Write and execute code for calculations and analysis - Return Markdown reports with inline citations and structured sources when literature evidence is relevant - Run asynchronously through polling or webhooks ## Public benchmark results ### Artificial Biological Intelligence Index beta - 19 AI models tested on 115 biology questions from LABBench2, HLE-Gold, and CGBench - Includes combined scores, estimated cost, and response time - Interactive results: https://sparkit.science/benchmarks/abi ### Research agent benchmarks Measured August 2026. Exact systems and test setups are named below; SPARKIT does not generalize these results to other products or model versions. - **HLE-Gold:** 149 questions; biology / medicine + chemistry - SPARKIT: 64.4% (SPARKIT research-agent run) - Claude Opus 5: 53.0% (Direct model call) - GPT-5.6-Sol: 39.0% (Direct model call) - Methods and limitations: https://sparkit.science/benchmarks/hle-gold Canonical benchmark index: https://sparkit.science/benchmarks ## Interfaces - API documentation: https://sparkit.science/docs - Python SDK: https://pypi.org/project/sparkit-science/ - MCP integration: https://sparkit.science/docs/integrations - Dashboard: https://app.sparkit.science/research ## Commercial information - Evaluation and integration inquiries: https://sparkit.science/contact - Research agent plans and terms: https://sparkit.science/pricing ## Important limitations SPARKIT is an LLM-driven system and can be wrong. Citations may be misattributed, sources may be summarized inaccurately, and conclusions may overstate the evidence. Verify cited sources before using an output for clinical, regulatory, legal, or other high-stakes decisions. ## Canonical pages - Services: https://sparkit.science/#services - Benchmarks: https://sparkit.science/benchmarks - ABI beta: https://sparkit.science/benchmarks/abi - Documentation: https://sparkit.science/docs - FAQ: https://sparkit.science/faq - Blog: https://sparkit.science/blog - About and authorship: https://sparkit.science/about - Full reference: https://sparkit.science/llms-full.txt