-
Anthropic Signs SpaceX Colossus Compute Deal, Doubles Claude Code Rate Limits
Anthropic
tools
-
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
Mind Lab
research
-
vLLM v0.25.0: Model Runner V2 Default, PagedAttention Retired, Transformers Backend Parity
tools
-
AMD and Anthropic announce 2-gigawatt Instinct MI450 compute deal with up to $5B AMD investment
Anthropic
industry
-
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
University of Illinois at Urbana-Champaign
research
-
Prime Intellect Releases prime-rl v0.6.0 for Agentic RL on Trillion-Parameter MoE Models
Prime Intellect
research
-
Google's Project Suncatcher to fly TPUs on SpaceX rideshare mission
Google DeepMind
research
-
LongLive-2.0: NVFP4 Parallel Infrastructure for Long Video Generation (NVIDIA, 1,220 HF upvotes)
NVIDIA
research
-
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
SLAI
research
-
Miles v0.1: a production-level RL post-training system
RadixArk
research
-
Yandex Open-Sources YaFF Data Format, Saving Up to 20% Server Capacity
Yandex
tools
-
Meta Announces "Meta Compute" Cloud Business to Monetize Surplus AI Infrastructure
Meta AI
industry
-
Meta raises 2026 AI capex guidance to $130-145B as free cash flow collapses
Meta
industry
-
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
University of Illinois Urbana-Champaign
research
-
Anthropic Signs $1.8 Billion Seven-Year Cloud Computing Deal with Akamai
Anthropic
industry
-
Anthropic signs $35 billion cloud deal with Nvidia-backed Lambda
Anthropic
industry
-
DeepSeek plans a 160,000-chip Huawei Ascend cluster in Inner Mongolia
DeepSeek
industry
-
ByteDance secures record $29.6 billion loan from ~30 banks for AI buildout
ByteDance
industry
-
Z.AI launches $5 billion Hong Kong raise: $2B share placement plus $3B convertible bonds
Zhipu AI (Z.AI)
industry
-
vLLM v0.21.0rc1: Python 3.14, CUDA 13.0, and Transformers v5 Compatibility
tools
-
ELDR: Expert-Locality-Aware Routing Cuts MoE Serving Latency by up to 14%
Microsoft Research
research
-
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Mind Lab
research
-
Dharma-AI lifts GPU-cluster utilization 53.6% → 87.0% by encoding physical constraints into the scheduler
Dharma-AI
tools
-
OpenAI details Habitat, the storage platform serving 70M requests per second
OpenAI
industry
-
Google advances Private AI Compute with secure, server-side memory
Google DeepMind
research
-
Meta and Anthropic reportedly in early talks over $10 billion compute lease deal
Anthropic
industry
-
VK Tech Reduces VK Data Platform Infrastructure Requirements 2.5× for AI Deployments
VK AI
tools