#benchmark
- Microsoft Build 2026: MAI Model Family Launched to Power GitHub Copilot Without OpenAI Dependency Microsoft models-llm
- xAI Releases Grok 4.3 with 1M Context, 40-60% Price Cuts, and Agentic Benchmark Gains xAI models-llm
- SenseNova-U1: Open-Source Unified Multimodal Understanding and Generation via NEO-unify SenseTime research
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Images Technion research
- EVA-Bench: End-to-End Framework for Evaluating Voice Agents ServiceNow AI research
- ExploitBench: Claude Mythos Preview and GPT-5.5 Develop Real Browser Exploits Autonomously Anthropic research
- VibeThinker-3B Reaches Frontier-Level Reasoning Benchmarks via Curriculum RL WeiboAI research
- Google DeepMind's AI Co-Mathematician Reaches 48% on FrontierMath Tier 4 Google DeepMind research
- Baidu Releases ERNIE 5.1 at 6% of Industry Pre-Training Cost, Enters Global Top-10 Search Baidu models-llm
- RubricEM: Meta-RL with Rubric-Guided Policy Decomposition Beyond Verifiable Rewards Google research
- SOOHAK: Frontier LLMs Solve Hard Math But Fail to Recognize Unsolvable Problems research
- ENPIRE: AI Coding Agents Close the Loop on Physical Robotics Research Without Human Intervention NVIDIA / Carnegie Mellon University / UC Berkeley research
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting Hao AI Lab, UC San Diego research
- OpenAI Releases GeneBench-Pro, a Frontier Benchmark for AI Agents in Biology OpenAI research
- LingBot-VLA 2.0: Bridging the Gap Between Foundation VLA Models and Real-World Deployment LingBot Team research
- ByteDance EdgeBench: Agent Learning Speed Doubles Every Three Months ByteDance research
- HumanCLAW: Can Vision-Language Models Act Through a Body? Meta Research research
- DeepSeek Puts V4-Flash-0731 API Into Public Beta, Beating Its Own Flagship on Agent Benchmarks DeepSeek models-llm
- MaxProof: MiniMax Model Exceeds IMO and USAMO Gold-Medal Thresholds on Formal Math MiniMax research
- Mistral Releases Leanstral 1.5: Open Formal-Verification Model for Lean 4 Mistral research
- LLM-as-a-Verifier: verification as an independent scaling axis for LLMs Stanford University / UC Berkeley / NVIDIA research
- AI Co-Mathematician: Google DeepMind Achieves 48% on FrontierMath Tier 4 Google DeepMind research
- MemLens: Benchmark for Multimodal Long-Term Memory in Vision-Language Models NVIDIA research
- Judge Circuits: Mechanistic Explanation of LLM-as-Judge Format Inconsistency research
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence (178 HF upvotes) Peking University / Shanghai Artificial Intelligence Laboratory research
- MMSkills: Reusable Multimodal Skills for General Visual Agents (105 HF upvotes) Shanghai Jiao Tong University research
- Crafter: Multi-Agent Harness for Editable Scientific Figure Generation Scores +16pt Over Baselines (103 HF Upvotes) Tsinghua University research
- EvoArena: LLM Agents Score Only 40% on Dynamic Evolving Environments MIT / NUS / Salesforce research
- WeaveBench: Computer-Use Agents Fail at Hybrid GUI+CLI Tasks — 41% Pass Rate Microsoft Research research
- Anthropic Study: Domain Expertise Drives Agentic Coding Success, Not Programming Background Anthropic research
- GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents research
- The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary research
- SingGuard: Runtime Policy-Adaptive Multimodal LLM Guardrail with 56K-Example Benchmark inclusionAI research
- PerceptionRubrics: Atomic Rubric Evaluation Reveals 8% Perception Gap Between Open and Closed Models research
- AutomationBench-AA: 657-task independent benchmark for AI agent SaaS automation Artificial Analysis tools
- Super Weights in LLMs: Why High-Salience Parameters Fail as Fine-Tuning Targets Amazon research
- ABot-AgentOS: General Robotic Agent OS with Lifelong Multi-modal Memory Alibaba research
- Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization research
- Anthropic's Project Pilot tests whether AI models can fly drones Anthropic research
- Programming with Data: test-driven data engineering for self-improving LLMs OpenDataLab research
- AutoResearchBench — a benchmark for autonomous scientific literature search by AI agents BAAI research
- GigaChat Passes Engineering Certification at Moscow Power Engineering Institute Sber industry
- Soohak: 64 Mathematicians Build Research-Level Benchmark That Stumps Frontier LLMs Seoul National University research
- EvoArena: LLM Agents Score Only 39.6% on Dynamic Evolving Environments Benchmark MIT research
- Are We Ready For an Agent-Native Memory System? SJTU Benchmarks 12 Architectures research
- AgenticSTS: Bounded-Memory Testbed for Long-Horizon LLM Agents Alaya Studio research
- EvoPolicyGym: Evaluating Iterative RL Policy Self-Improvement by Coding Agents University of Macau / CUHK research
- SWE-Together: Multi-Turn Benchmark for Coding Agent Evaluation research
- Ideas Have Genomes: Frontier LLMs Score Only 27% on Scientific Lineage Reasoning Shanghai Jiao Tong University research
- Executable World Models for ARC-AGI-3: Coding-Agent Approach Without Game-Specific Logic research
- Learning, Fast and Slow: Dual-Weight Architecture for Continual LLM Adaptation research
- SubtleMemory: Benchmark Reveals Agents Systematically Fail Fine-Grained Relational Memory research
- VideoKR: 315K-Example Training Corpus for Knowledge- and Reasoning-Intensive Video Understanding Yale University research
- SWE-Explore: Benchmarking Repository Exploration as the Binding Constraint in Coding Agents Shanghai Jiao Tong University research
- StylisticBias: 15 Visual Attributes Account for 80% of Social Bias in Multimodal LLMs research
- Will Scaling Improve Social Simulation with LLMs? A Study of 85 Models Stanford / Columbia / Tsinghua research
- UniClawBench: Benchmark for Proactive AI Agents on Real-World Tasks University of Hong Kong research
- AdvancedMathBench: Benchmark Suite for Advanced Mathematical Proof Generation and Verification InternLM research
- GuardianAgentBench: Where Agents Fail and How to Guard Them research
- LLMs Get Lost in Evolving User Intent Microsoft Research research