-
Google DeepMind Invests $75M in A24, Forms First AI Research Partnership with a Film Studio
Google DeepMind
industry
-
Anthropic Launches Claude Science, an AI Workbench for Researchers
Anthropic
tools
-
OpenAI Releases GeneBench-Pro, a Frontier Benchmark for AI Agents in Biology
OpenAI
research
-
RoPE Provably Fails at Long Contexts: Locality Bias and Token Consistency Both Break
research
-
MiniMax Sparse Attention: 28× Compute Reduction at 1M-Token Context with No Quality Loss
MiniMax
research
-
MaxProof: MiniMax Model Exceeds IMO and USAMO Gold-Medal Thresholds on Formal Math
MiniMax
research
-
Google DeepMind Publishes AlphaEvolve One-Year Impact Report
Google DeepMind
research
-
Crafter: Multi-Agent Harness for Editable Scientific Figure Generation Scores +16pt Over Baselines (103 HF Upvotes)
Tsinghua University
research
-
GrepSeek: Training Search Agents for Direct Corpus Interaction via Shell Commands (93 HF Upvotes)
University of Massachusetts Amherst
research
-
EvoArena: LLM Agents Score Only 40% on Dynamic Evolving Environments
MIT / NUS / Salesforce
research
-
WeaveBench: Computer-Use Agents Fail at Hybrid GUI+CLI Tasks — 41% Pass Rate
Microsoft Research
research
-
InterleaveThinker: RL Planner+Critic Pipeline for Interleaved Text-and-Image Generation
CUHK Multimedia Lab
research
-
Metis: Memory Foundation Model
MemTensor, Renmin University, NUS, Shanghai Jiao Tong University, Tongji University
research
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Horizon Research, Frontis.AI, Tsinghua University
research
-
PhiZero: A World Model Built Around Physical Language
NLPR, Institute of Automation, Chinese Academy of Sciences
research
-
OpenAI Launches Economic Research Exchange for AI Impact Studies
OpenAI
industry
-
VK Research: classical recommendation algorithms can account for users' future interests
VK
research
-
BetaPRM: Uncertainty-Aware Process Rewards Cut Reasoning Token Use by 33%
research
-
Google DeepMind and Partners Launch $10M Multi-Agent AI Safety Research Fund
Google DeepMind
industry
-
Anthropic Publishes First Public Record: 52,000-Person Survey on US AI Attitudes
Anthropic
research
-
Anthropic launches Economic Index connector for Claude
Anthropic
tools
-
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
research
-
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Stanford University
research
-
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
research
-
QQWorld: Quantile-Quantile Matching for World Model Regularization
research
-
Google DeepMind Takes Minority Stake in CCP Games for Multi-Agent Research in EVE Online
Google DeepMind
industry