-
DeepSeek V4: official open-source release with Day-0 adaptation for Huawei Ascend
DeepSeek
models-llm
-
xAI completes Grok 4.3 API rollout with 1M context, native video, and ~40% price cut
xAI
models-llm
-
MiniMax Releases M3: Open-Weight Frontier Model with 1M-Token Context and MSA Architecture
MiniMax
models-llm
-
NVIDIA Nemotron 3 Ultra: Open 550B MoE Model Now Available for Agentic Workloads
NVIDIA
models-llm
-
MiniMax M3 Open Weights Released: 1M Context, MoE, Frontier Coding
MiniMax
models-llm
-
Zhipu AI Open-Sources GLM-5.2 Under MIT License with 1M Token Context
Zhipu AI
models-llm
-
Zhipu AI Releases GLM-5.2 Open Weights: 753B MoE with 1M-Token Context under MIT License
Zhipu AI / Z.ai
models-llm
-
Moonshot AI launches Kimi K3, a 2.8T-parameter open-weight model
Moonshot AI
models-llm
-
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
Apodex
research
-
DeepSeek V4.1 Flash goes GA with open MIT weights and V4 Pro retirement
DeepSeek
models-llm
-
Tencent open-sources Hy4 preview: 770B / 49B-active MoE with Gated DSA + iHC, 1M context, Apache 2.0
Tencent
models-llm
-
SU-01: Gold-Medal-Level Olympiad Reasoning via Curriculum SFT and Two-Stage RL
SU-01 Team
research
-
A Systematic Analysis of Hybrid Linear Attention: 72-Model Study
ByteDance Seed
research
-
DeepSeek-V4-Pro reaches general availability, prices jump up to 1,100% from Aug 16
DeepSeek
models-llm
-
Tencent Hunyuan open-sources SAS sparse-attention routers trained end-to-end on Qwen3
Tencent Hunyuan
research
-
RoPE Provably Fails at Long Contexts: Locality Bias and Token Consistency Both Break
research
-
Moonshot AI Releases Kimi K2.7-Code: 1T-Parameter Open-Weight Coding Model with Vision
Moonshot AI
models-llm
-
MiniMax Sparse Attention: 28× Compute Reduction at 1M-Token Context with No Quality Loss
MiniMax
research
-
Kimi K3 Technical Report: Kimi Delta Attention and Stable LatentMoE Architecture Detailed
Moonshot AI
research
-
IBM releases Granite 4.2 reasoning LLMs (3B/8B/30B, Apache 2.0, 512K context) trained on ~15T tokens with GB200 NVL72 GRPO RL
IBM
models-llm
-
MemLens: Benchmark for Multimodal Long-Term Memory in Vision-Language Models
NVIDIA
research
-
Echo-Infinity: Real-Time Infinite Video Generation via Learnable Memory Query
research
-
GitHub Copilot Gets 1M Token Context Window and Configurable Reasoning Levels
GitHub / Microsoft
tools
-
vLLM Adds Day-0 Support for MiniMax M3 Open Weights with 1M-Context Sparse Attention
MiniMax
tools
-
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
Zhipu AI / Tsinghua University
research
-
Metis: Memory Foundation Model
MemTensor, Renmin University, NUS, Shanghai Jiao Tong University, Tongji University
research
-
Demystifying Agent Skills: Why They Work — Until They Don't
UC San Diego (Zhiyuan Jiang, Mengdi Wang, Yijiang Li et al.)
research
-
Zhipu releases GLM-5.3-Flash, first natively multimodal GLM-5 model
Zhipu AI / Z.ai
models-llm
-
Qwen3.8-Flash-Next released as experimental Qwen4 architecture preview
Alibaba (Qwen Team)
models-llm
-
Qwen team details the Qwen3.8-Next architecture: hybrid attention, n-gram embeddings, Muon
Qwen (Alibaba)
research
-
LatentPress: compressing context into continuous memory tokens a frozen LLM reads directly
research
-
Ant Group open-sources Ling-3.0-flash-Fin, its first finance-enhanced model
Ant Group
models-llm
-
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek
research
-
Do Language Models Need Sleep? Offline Recurrence as Memory Consolidation for Improved Inference
Google / CMU
research
-
Zhipu AI Releases GLM-5.2: 744B MoE with 1M-Token Context and Coding-First Design
Zhipu AI
models-llm
-
AgenticSTS: Bounded-Memory Testbed for Long-Horizon LLM Agents
Alaya Studio
research
-
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
MIT / NVIDIA
research
-
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Shanghai AI Lab
research
-
Prime Agent: A Self-Improving RLM Harness
Prime Intellect
research
-
ReWorld: An Interactive World Model with Long-Horizon Memory
TongyiLab (Alibaba)
research
-
SubtleMemory: Benchmark Reveals Agents Systematically Fail Fine-Grained Relational Memory
research
-
SearchSwarm: Delegation Intelligence for LLM Agents in Long-Horizon Deep Research
research
-
FlashMorph: Data-Driven Hybrid Attention Layer Placement via Learnable Gates
ByteDance Seed
research
-
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Mind Lab
research
-
vLLM previews production-scale support for Moonshot AI's Kimi K3 ahead of its weight release
vLLM / Moonshot AI
tools
-
Motif 3: Technical Report
Motif Technologies
research
-
llama.cpp Sep 1-2 wave: +4.9% generation from n-gram lookup fix, fused CUDA MoE reduction
ggml
tools
-
Language Models Can Control Their Own Attention
research