#evals
- GigaChat 3.5 Ultra passes professional-retraining information security exam, scoring 14% above pass threshold Sber research
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks research
- Cline open-sources Terminal-Bench-based evals for open-weight coding agents Cline tools
- Demystifying Agent Skills: Why They Work — Until They Don't UC San Diego (Zhiyuan Jiang, Mengdi Wang, Yijiang Li et al.) research
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis University of Science and Technology of China research
- Anthropic commits $5M to independent research on AI's impact on wellbeing Anthropic research
- Anthropic opens aggregate Claude usage data to external researchers via Anthropic Insights pilot (Stanford, Oxford, METR) Anthropic research
- Anthropic publishes alignment assessment of Claude incidents in cybersecurity evaluations Anthropic research
- Anthropic Frontier Red Team measures intelligence-targeting and weapons capabilities Anthropic research
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? OpenMOSS research
- Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains research
- ASI-Bench: At the Dawn of Artificial Superintelligence 42-author consortium (lead: Junwei Zhou; incl. Chi Wang, Yilun Hao, Yuantao Zhai) research
- Phoenix v20.4.0 ships in-process MCP toolset and retrieval-relevance evaluator Arize AI tools