#benchmarks
- LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers University of Illinois at Urbana-Champaign research
- Long-Horizon-Terminal-Bench: Testing Agent Limits on Long-Horizon Terminal Tasks Tencent Hunyuan research
- Demystifying Agent Skills: Why They Work — Until They Don't UC San Diego (Zhiyuan Jiang, Mengdi Wang, Yijiang Li et al.) research
- EnvHarness: Awakening Static Worlds for Agent Learning Google research
- K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs research
- VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Tencent research
- HarnessEval-W: Agentifying the Evaluation of Visual Worlds NTU / MirroS-Lab consortium research
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? OpenMOSS research
- IdeaGene-Bench: Benchmarking Scientific Lineage Reasoning and Idea Generation Shanghai Jiao Tong University research
- ASI-Bench: At the Dawn of Artificial Superintelligence 42-author consortium (lead: Junwei Zhou; incl. Chi Wang, Yilun Hao, Yuantao Zhai) research
- MWS AI Releases Cotype Pro 3 (27B) and Light 3 (9B) for Enterprise AI Agents MWS AI models-llm