ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage

Research official 3 src. ~1 min

ClawProBench is a trace-aware benchmark for stateful-runtime AI agents, instantiated on OpenClaw. It pairs a 102-scenario full profile with a 68-scenario frozen holdout and uses a safety-gated scoring formula derived from execution traces, evaluating 68 configurations on the full profile and 37 on the holdout.

Why it matters

HF Daily Papers Aug 25 at 822 upvotes — one of the highest-voted benchmark releases of the window. Highlights that correctness-only rankings diverge from process-aware rankings (Spearman 0.13 full vs holdout), arguing for trace-level evaluation of agents.

Importance: 4/5

HF Daily 822 upvotes

Sources