Benchmarks, evaluations, datasets, and technical research tracking the Claw category.
4 resources
An end-to-end benchmark for autonomous scientific research agents, covering real datasets, code, figures, reports, and peer-review-style evaluation across multiple disciplines.
A benchmark for always-on personal assistants that measures long-horizon event understanding, interconnected digital services, and cross-device GUI and CLI interaction.
An open-source toolkit for generating, executing, and grading environments for Claw-like agents, paired with the Auto-ClawEval benchmark of 1,040 environments across 24 categories.
A benchmark for comparing heterogeneous agent harnesses—described by the authors as claws—on coding tasks under controlled settings.