#architecture
- MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling Shanghai AI Lab research
- Mean Mode Screaming: Training Pathology Fix Enables 1000-Layer Diffusion Transformers research
- Lance: 3B Unified Multimodal Model for Understanding, Generation, and Editing (314 HF upvotes) ByteDance Research research
- A Systematic Analysis of Hybrid Linear Attention: 72-Model Study ByteDance Seed research
- Metis: Memory Foundation Model MemTensor research
- Echo-Infinity: Real-Time Infinite Video Generation via Learnable Memory Query research
- Hidden Decoding at Scale: A New Axis for LLM Capacity Without Backbone Growth research
- Metis: Memory Foundation Model MemTensor, Renmin University, NUS, Shanghai Jiao Tong University, Tongji University research
- Do Language Models Need Sleep? Offline Recurrence as Memory Consolidation for Improved Inference Google / CMU research
- Wan-Streamer v0.1: End-to-End Real-Time Interactive Foundation Model Under 550ms Latency Wan-AI research
- Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE MIT / NVIDIA research
- Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning Shanghai AI Lab research
- Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and a Fix research
- Cola DLM: Continuous Latent Diffusion Language Model with Competitive Scaling research
- FlashMorph: Data-Driven Hybrid Attention Layer Placement via Learnable Gates ByteDance Seed research
- xHC: Expanded Hyper-Connections scale residual-stream parallelism past prior limits Shanghai Jiao Tong University / Xiaohongshu / USTC / Peking University / CUHK research
- Motif 3: Technical Report Motif Technologies research