#efficiency
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM research
- Qwen-Image-2.0: Unified Image Generation and Editing at 2K Resolution, Top-1 on AI Arena Alibaba research
- Baidu Releases ERNIE 5.1 at 6% of Industry Pre-Training Cost, Enters Global Top-10 Search Baidu models-llm
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting Hao AI Lab, UC San Diego research
- Dockerless: Environment-Free Program Verifier for Coding Agents ByteDance research
- DOPD: Dual On-Policy Distillation with Advantage-Aware Token Routing research
- A Systematic Analysis of Hybrid Linear Attention: 72-Model Study ByteDance Seed research
- Program-as-Weights: Compiling Task Specs into LoRA Adapters for On-Device Inference University of Waterloo / Harvard research
- MrFlow: Training-Free 10-25x Speedup for Flow-Matching Text-to-Image Models Beihang University research
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Microsoft research
- Kwai Keye-VL-2.0: Open-Source 30B MoE Multimodal Model with 256K Context for Long Video Kwai research
- MiniMax Sparse Attention: 28× Compute Reduction at 1M-Token Context with No Quality Loss MiniMax research
- Moebius: 0.2B Lightweight Image Inpainting Framework Matches 11.9B FLUX Model Huazhong University of Science and Technology research
- BDH-CQ: Recurrent Latent Reasoning Model Breaks ARC-AGI-1 Cost-Efficiency Frontier Pathway research
- Orthrus: 7.8x Inference Speedup for Qwen3 via Autoregressive-Diffusion KV Sharing research
- SANA-WM: Minute-Scale 720p World Modeling on a Single GPU NVIDIA research
- ThoughtFold: Introspective Preference Learning Cuts Reasoning Tokens by 56% Without Accuracy Loss research
- FastContext: Specialized Exploration Subagent Cuts Coding Agent Token Usage by 60% Microsoft / Shanghai Jiao Tong University research
- Quantized Reasoning Models Think They Need to Think Longer, but They Do Not Meta research
- vLLM v0.24.0: Model Runner V2 Default, Rust Frontend, SM90 FP8 Speedups vLLM tools
- Program-as-Weights: Compile-Once Adapter Paradigm Matches 32B Models at 1/50 the Memory research
- Super Weights in LLMs: Why High-Salience Parameters Fail as Fine-Tuning Targets Amazon research
- Are We Ready For an Agent-Native Memory System? SJTU Benchmarks 12 Architectures research
- BlockPilot: Instance-Adaptive Block Size for Diffusion-Based Speculative Decoding research
- Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE MIT / NVIDIA research
- Scalable Visual Pretraining for Language Intelligence research
- On the Geometry of On-Policy Distillation: A Training Paradigm Distinct from SFT and RLVR Hong Kong University of Science and Technology research
- SHERLOC: Structured Diagnostic Localization Cuts Code Repair Token Usage by 36.7% research
- OPRD: On-Policy Representation Distillation for Post-Training LLMs research
- ELDR: Expert-Locality-Aware Routing Cuts MoE Serving Latency by up to 14% Microsoft Research research
- FlashMorph: Data-Driven Hybrid Attention Layer Placement via Learnable Gates ByteDance Seed research
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Microsoft research
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation research