-
DeepSeek V4: official open-source release with Day-0 adaptation for Huawei Ascend
DeepSeek
models-llm
-
Tencent Officially Releases Hunyuan Hy3: 295B MoE Model with Agent and Reasoning Capabilities
Tencent
models-llm
-
Moonshot Kimi K3 Open-Weight Release: 2.8T Parameter MoE Model on Hugging Face
Moonshot AI
models-llm
-
MiniMax Releases M3: Open-Weight Frontier Model with 1M-Token Context and MSA Architecture
MiniMax
models-llm
-
NVIDIA Nemotron 3 Ultra: Open 550B MoE Model Now Available for Agentic Workloads
NVIDIA
models-llm
-
MiniMax M3 Open Weights Released: 1M Context, MoE, Frontier Coding
MiniMax
models-llm
-
Zhipu AI Open-Sources GLM-5.2 Under MIT License with 1M Token Context
Zhipu AI
models-llm
-
Zhipu AI Releases GLM-5.2 Open Weights: 753B MoE with 1M-Token Context under MIT License
Zhipu AI / Z.ai
models-llm
-
ByteDance Launches Doubao-Seed-2.1 Pro Flagship LLM at FORCE Conference
ByteDance / Doubao
models-llm
-
Sber Releases GigaChat 3.5 Ultra: 432B Open-Source MoE Flagship
Sber
models-llm
-
Tencent Releases Hunyuan Hy3: 295B Open-Weight MoE Model Under Apache 2.0
Tencent
models-llm
-
Thinking Machines Lab Releases Inkling: 975B Open-Weight Multimodal MoE
Thinking Machines Lab
models-llm
-
Moonshot AI launches Kimi K3, a 2.8T-parameter open-weight model
Moonshot AI
models-llm
-
Zyphra Releases ZAYA1-8B: Open Reasoning MoE Model Trained on AMD Hardware
Zyphra
models-llm
-
Lance: 3B Unified Multimodal Model for Understanding, Generation, and Editing (314 HF upvotes)
ByteDance Research
research
-
JetBrains Open-Sources Mellum2: 12B MoE Coding Model for Multi-Model Pipelines
JetBrains
models-llm
-
Cohere North Mini Code: 30B Apache-2.0 MoE Coding Model for Agentic Workflows
Cohere
models-llm
-
Kimi K2.7-Code HighSpeed: 6× Throughput for Production Coding Agent Pipelines
Moonshot AI
models-llm
-
Prime Intellect Releases prime-rl v0.6.0 for Agentic RL on Trillion-Parameter MoE Models
Prime Intellect
research
-
DeepReinforce Releases Ornith-1.0: Open-Source Coding Models That Learn Their Own RL Scaffolds
DeepReinforce
tools
-
DeepSeek V4 Stable Release Set for Mid-July with First Time-of-Day API Pricing
DeepSeek
models-llm
-
DeepSeek Confirms V4 Official Launch for Mid-July with Peak-Time API Pricing
DeepSeek
models-llm
-
NVIDIA Releases Nemotron-Labs-Audex-30B-A3B: Unified Audio-Text MoE Model
NVIDIA
audio
-
DeepSeek V4 Graduates from Preview to General Availability with Peak-Hour API Pricing
DeepSeek
models-llm
-
Alibaba previews Qwen3.8-Max, a 2.4-trillion-parameter multimodal flagship
Alibaba / Qwen
models-llm
-
Kwai Keye-VL-2.0: Open-Source 30B MoE Multimodal Model with 256K Context for Long Video
Kwai
research
-
Moonshot AI Releases Kimi K2.7-Code: 1T-Parameter Open-Weight Coding Model with Vision
Moonshot AI
models-llm
-
Ring-Zero: Scaling Zero RL to 1 Trillion Parameters with Emergent Reasoning Behaviors
Ant Group
research
-
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
SLAI
research
-
Kimi K3 Technical Report: Kimi Delta Attention and Stable LatentMoE Architecture Detailed
Moonshot AI
research
-
vLLM Adds Day-0 Support for MiniMax M3 Open Weights with 1M-Context Sparse Attention
MiniMax
tools
-
Kimi K2.7 Code (1T params, 32B active) expands to GitHub Copilot Business and Enterprise
Moonshot AI
tools
-
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
SLAI
research
-
Zhipu AI Releases GLM-5.2: 744B MoE with 1M-Token Context and Coding-First Design
Zhipu AI
models-llm
-
GLM-5.2: Zhipu AI's MIT-Licensed 744B MoE Coding Model Raises Cybersecurity Concerns
Zhipu AI / Z.ai
models-llm
-
Alibaba Makes Qwen3.8-Max Widely Accessible via API, Launches QwenWork Public Beta
Alibaba / Qwen
models-llm
-
Sber unveils Kandinsky 6.0 Image — flagship image generation model
Sber
image
-
ELDR: Expert-Locality-Aware Routing Cuts MoE Serving Latency by up to 14%
Microsoft Research
research
-
llama.cpp Adds Tencent Hunyuan 3, Minimax2 Eagle3 Speculative Decoding, and SYCL Fused MoE
tools
-
xHC: Expanded Hyper-Connections scale residual-stream parallelism past prior limits
Shanghai Jiao Tong University / Xiaohongshu / USTC / Peking University / CUHK
research