-
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
research
-
Qwen-Image-2.0: Unified Image Generation and Editing at 2K Resolution, Top-1 on AI Arena
Alibaba
research
-
Baidu Releases ERNIE 5.1 at 6% of Industry Pre-Training Cost, Enters Global Top-10 Search
Baidu
models-llm
-
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
Hao AI Lab, UC San Diego
research
-
Dockerless: Environment-Free Program Verifier for Coding Agents
ByteDance
research
-
DOPD: Dual On-Policy Distillation with Advantage-Aware Token Routing
research
-
A Systematic Analysis of Hybrid Linear Attention: 72-Model Study
ByteDance Seed
research
-
Program-as-Weights: Compiling Task Specs into LoRA Adapters for On-Device Inference
University of Waterloo / Harvard
research
-
MrFlow: Training-Free 10-25x Speedup for Flow-Matching Text-to-Image Models
Beihang University
research
-
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Microsoft
research
-
Tencent Hunyuan open-sources SAS sparse-attention routers trained end-to-end on Qwen3
Tencent Hunyuan
research
-
Kwai Keye-VL-2.0: Open-Source 30B MoE Multimodal Model with 256K Context for Long Video
Kwai
research
-
MiniMax Sparse Attention: 28× Compute Reduction at 1M-Token Context with No Quality Loss
MiniMax
research
-
Moebius: 0.2B Lightweight Image Inpainting Framework Matches 11.9B FLUX Model
Huazhong University of Science and Technology
research
-
BDH-CQ: Recurrent Latent Reasoning Model Breaks ARC-AGI-1 Cost-Efficiency Frontier
Pathway
research
-
Uno: lossless LLM speedups via diffusion-augmented autoregression
MBZUAI
research
-
WMRL: scaling automatic research agents via world models
University of Illinois / NVIDIA
research
-
NCP-ArchPreview: latent-space language models at 8.9B via Next Concept Prediction
Shanghai AI Lab
research
-
Vidu S2: real-time interactive, editable, and spatial video generation
ShengShu Intelligence
research
-
ZGCM-1: a fully open, extremely efficient 7B foundation model for math and agentic search
ZGCAGI
research
-
Dream-RSI: recursive self-improvement through evolving worlds
Google
research
-
Orthrus: 7.8x Inference Speedup for Qwen3 via Autoregressive-Diffusion KV Sharing
research
-
SANA-WM: Minute-Scale 720p World Modeling on a Single GPU
NVIDIA
research
-
ThoughtFold: Introspective Preference Learning Cuts Reasoning Tokens by 56% Without Accuracy Loss
research
-
FastContext: Specialized Exploration Subagent Cuts Coding Agent Token Usage by 60%
Microsoft / Shanghai Jiao Tong University
research
-
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
Meta
research
-
vLLM v0.24.0: Model Runner V2 Default, Rust Frontend, SM90 FP8 Speedups
vLLM
tools
-
Program-as-Weights: Compile-Once Adapter Paradigm Matches 32B Models at 1/50 the Memory
research
-
Super Weights in LLMs: Why High-Salience Parameters Fail as Fine-Tuning Targets
Amazon
research
-
Qwen team details the Qwen3.8-Next architecture: hybrid attention, n-gram embeddings, Muon
Qwen (Alibaba)
research
-
Google DeepMind ships agentic video understanding in Gemini
Google DeepMind
tools
-
LatentPress: compressing context into continuous memory tokens a frozen LLM reads directly
research
-
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek
research
-
Are We Ready For an Agent-Native Memory System? SJTU Benchmarks 12 Architectures
research
-
BlockPilot: Instance-Adaptive Block Size for Diffusion-Based Speculative Decoding
research
-
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
MIT / NVIDIA
research
-
Scalable Visual Pretraining for Language Intelligence
research
-
SMELT: looped MoE transformers match baseline scaling at matched compute
ByteDance Seed
research
-
COBRA-Skills: contextual bandit-guided evolution for agent skills
Chinese University of Hong Kong, Shenzhen
research
-
On the Geometry of On-Policy Distillation: A Training Paradigm Distinct from SFT and RLVR
Hong Kong University of Science and Technology
research
-
SHERLOC: Structured Diagnostic Localization Cuts Code Repair Token Usage by 36.7%
research
-
OPRD: On-Policy Representation Distillation for Post-Training LLMs
research
-
ELDR: Expert-Locality-Aware Routing Cuts MoE Serving Latency by up to 14%
Microsoft Research
research
-
FlashMorph: Data-Driven Hybrid Attention Layer Placement via Learnable Gates
ByteDance Seed
research
-
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Microsoft
research
-
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
research
-
Does on-policy distillation really distill? Teacher-free OPSA beats it on AIME24
Purdue University
research
-
Language Models Can Control Their Own Attention
research
-
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Cerebras
research
-
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
research