Decision information
Verified: 8/14/2026- Ecosystem category
- Research & Benchmarks
- Provider
- InternScience
- Content type
- Research
- Pricing
- Open Source
- Open source
- Yes
- License
- MIT
- Maintenance
- Active
- Install methods
- Source
- Platforms
- Linux
- Interfaces
- Web, Terminal
Security signals (informational, not a security certification)
Source availableRequires shellRequires credentials
Verification source: https://arxiv.org/abs/2606.07591 · arXiv:2606.07591
Overview
An end-to-end benchmark for autonomous scientific research agents, covering real datasets, code, figures, reports, and peer-review-style evaluation across multiple disciplines.
What it provides
- Forty real-science tasks spanning ten disciplines
- Agent-agnostic execution with multimodal rubric scoring
- Supports OpenClaw, ResearchClaw, coding agents, and custom harnesses
AllClaw classification
This project supports the Claw ecosystem as research benchmarks; it is not presented as a standalone Claw agent.
Provider: InternScience. Primary source: https://arxiv.org/abs/2606.07591.
The classification and core claims were checked against the linked primary source on August 14, 2026.
Tags
benchmarkscientific-researchevaluationmulti-agentleaderboard