Back to Ecosystem
ClawEnvKit & Auto-ClawEval logo
Research & BenchmarksVerified

ClawEnvKit & Auto-ClawEval

An open-source toolkit for generating, executing, and grading environments for Claw-like agents, paired with the Auto-ClawEval benchmark of 1,040 environments across 24 categories.

Decision information

Verified: 8/14/2026
Ecosystem category
Research & Benchmarks
Provider
AIcell
Content type
Research
Pricing
Open Source
Open source
Yes
License
MIT
Maintenance
Active
Install methods
pip, Docker, Source
Platforms
macOS, Linux, Docker
Interfaces
Terminal, MCP, OpenClaw plugin

Security signals (informational, not a security certification)

Source availableRequires shellRequires credentials

Verification source: https://arxiv.org/abs/2604.18543 · arXiv:2604.18543

Overview

An open-source toolkit for generating, executing, and grading environments for Claw-like agents, paired with the Auto-ClawEval benchmark of 1,040 environments across 24 categories.

What it provides

  • Generates agent environments from natural-language specifications
  • Uses auditable mock services and deterministic verification
  • Evaluates multiple Claw harnesses through plugin, MCP, and shell adapters

AllClaw classification

This project supports the Claw ecosystem as research benchmarks; it is not presented as a standalone Claw agent.

Provider: AIcell. Primary source: https://arxiv.org/abs/2604.18543.

The classification and core claims were checked against the linked primary source on August 14, 2026.

Tags

benchmarkevaluationenvironment-generationdatasetagent-harness