ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage
ClawProBench is a trace-aware benchmark for stateful-runtime AI agents, instantiated on OpenClaw. It pairs a 102-scenario full profile with a 68-scenario frozen holdout and uses a safety-gated scoring formula derived from execution traces, evaluating 68 configurations on the full profile and 37 on the holdout.
Why it matters
HF Daily Papers Aug 25 at 822 upvotes — one of the highest-voted benchmark releases of the window. Highlights that correctness-only rankings diverge from process-aware rankings (Spearman 0.13 full vs holdout), arguing for trace-level evaluation of agents.
Importance: 4/5
HF Daily 822 upvotes
Sources
official
ClawProBench (arXiv 2608.22510)
official
HF Daily Papers: ClawProBench
official
ClawProBench reference harness