Benchmark Radar: a living database and search engine for AI benchmarks
Carnegie Mellon University
A continuously updated catalog and search engine for AI benchmarks spanning LLM, agentic, coding, reasoning, and safety evaluation, aggregating daily from 37 sources into 1,283 records and 12,916 score observations with preserved citation trails. It also ships analyses of benchmark saturation, adoption trends, and a prior-art search workflow for designing new evaluations.
Why it matters
Addresses the chronic inability of the field to find, version, and compare benchmarks; saturation and trend views make evaluation-fragmentation visible in one place.
Importance: 2/5
Notable evaluation-infrastructure paper, 80 upvotes on HF Daily