#attention
- A Systematic Analysis of Hybrid Linear Attention: 72-Model Study ByteDance Seed research
- MiniMax Sparse Attention: 28× Compute Reduction at 1M-Token Context with No Quality Loss MiniMax research
- FlashMorph: Data-Driven Hybrid Attention Layer Placement via Learnable Gates ByteDance Seed research