#benchmarks
- OpenAI launches GPT-6 Astra, its most intelligent and aligned model OpenAI models-llm
- LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers University of Illinois at Urbana-Champaign research
- ARC Prize shows GPT-6 Astra's 99.9% ARC-AGI-3 score came from OpenAI's custom harness, not the model OpenAI research
- Long-Horizon-Terminal-Bench: Testing Agent Limits on Long-Horizon Terminal Tasks Tencent Hunyuan research
- Specific releases Real-SWE benchmark built from private enterprise codebases Specific tools
- The Tasteful Agent: measuring 'taste' in long-horizon tasks Microsoft research
- Demystifying Agent Skills: Why They Work — Until They Don't UC San Diego (Zhiyuan Jiang, Mengdi Wang, Yijiang Li et al.) research
- EnvHarness: Awakening Static Worlds for Agent Learning Google research
- K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs research
- VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Tencent research
- HarnessEval-W: Agentifying the Evaluation of Visual Worlds NTU / MirroS-Lab consortium research
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? OpenMOSS research
- Benchmark Radar: a living database and search engine for AI benchmarks Carnegie Mellon University research
- RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Qwen (Alibaba) research
- IdeaGene-Bench: Benchmarking Scientific Lineage Reasoning and Idea Generation Shanghai Jiao Tong University research
- ASI-Bench: At the Dawn of Artificial Superintelligence 42-author consortium (lead: Junwei Zhou; incl. Chi Wang, Yilun Hao, Yuantao Zhai) research
- Last Translation Benchmark: human-authored tasks that still break frontier translation models research
- SpatialBlock: Teaching LVLMs Spatial Reasoning with Synthetic Block-Stacking KAIST AI research
- MWS AI Releases Cotype Pro 3 (27B) and Light 3 (9B) for Enterprise AI Agents MWS AI models-llm