#transformers
- Your Transformer Can Hold Two Thoughts at Once: evidence of linear superposition in LLMs research
- NCP-ArchPreview: latent-space language models at 8.9B via Next Concept Prediction Shanghai AI Lab research
- Hugging Face Transformers: Async Continuous Batching Achieves 22% Inference Speedup Hugging Face tools
- Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and a Fix research
- Agentic Transformers Provably Learn to Search via Reinforcement Learning research
- Transformers v5.16.0 lands Qwen4-Exp hybrid attention, ESMC + ESMFold2, GLM 5.3 Flash support in v5.16.1 Hugging Face tools