-
MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling
Shanghai AI Lab
research
-
Mean Mode Screaming: Training Pathology Fix Enables 1000-Layer Diffusion Transformers
research
-
Lance: 3B Unified Multimodal Model for Understanding, Generation, and Editing (314 HF upvotes)
ByteDance Research
research
-
A Systematic Analysis of Hybrid Linear Attention: 72-Model Study
ByteDance Seed
research
-
Metis: Memory Foundation Model
MemTensor
research
-
Recurrent Looped Transformer: viral technical report claims 'infinite temporal depth' for latent reasoning
Princeton University
research
-
NCP-ArchPreview: latent-space language models at 8.9B via Next Concept Prediction
Shanghai AI Lab
research
-
Echo-Infinity: Real-Time Infinite Video Generation via Learnable Memory Query
research
-
Hidden Decoding at Scale: A New Axis for LLM Capacity Without Backbone Growth
research
-
Metis: Memory Foundation Model
MemTensor, Renmin University, NUS, Shanghai Jiao Tong University, Tongji University
research
-
Qwen team details the Qwen3.8-Next architecture: hybrid attention, n-gram embeddings, Muon
Qwen (Alibaba)
research
-
Do Language Models Need Sleep? Offline Recurrence as Memory Consolidation for Improved Inference
Google / CMU
research
-
Wan-Streamer v0.1: End-to-End Real-Time Interactive Foundation Model Under 550ms Latency
Wan-AI
research
-
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
MIT / NVIDIA
research
-
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Shanghai AI Lab
research
-
SMELT: looped MoE transformers match baseline scaling at matched compute
ByteDance Seed
research
-
Infinite-Parameter LLMs: arXiv paper proposes generating weights from live data
Independent researchers
research
-
Language Model 'Shape': designing architectures around agent workflows
Alex Zhang (independent)
research
-
Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and a Fix
research
-
Cola DLM: Continuous Latent Diffusion Language Model with Competitive Scaling
research
-
FlashMorph: Data-Driven Hybrid Attention Layer Placement via Learnable Gates
ByteDance Seed
research
-
xHC: Expanded Hyper-Connections scale residual-stream parallelism past prior limits
Shanghai Jiao Tong University / Xiaohongshu / USTC / Peking University / CUHK
research
-
Motif 3: Technical Report
Motif Technologies
research