Benchmark Radar: a living database and search engine for AI benchmarks

Carnegie Mellon University

Research official + media 2 src. ~1 min

A continuously updated catalog and search engine for AI benchmarks spanning LLM, agentic, coding, reasoning, and safety evaluation, aggregating daily from 37 sources into 1,283 records and 12,916 score observations with preserved citation trails. It also ships analyses of benchmark saturation, adoption trends, and a prior-art search workflow for designing new evaluations.

Why it matters

Addresses the chronic inability of the field to find, version, and compare benchmarks; saturation and trend views make evaluation-fragmentation visible in one place.

Importance: 2/5

Notable evaluation-infrastructure paper, 80 upvotes on HF Daily

Sources