Your Transformer Can Hold Two Thoughts at Once: evidence of linear superposition in LLMs
Proposes the Superposition Linearity Hypothesis: mixing two text streams as inputs yields an output distribution approximately equal to the blend of each stream's separate predictions. The trait diminishes during pretraining but is recoverable with light fine-tuning, and the authors demonstrate dual-stream guided decoding from a single forward pass.
Why it matters
Second-most-upvoted HF paper for Sep 25 (55 upvotes); a testable empirical account of superposition directly in activations rather than in SAE features.
Importance: 3/5
Highly-upvoted HF Daily paper with a new testable interpretability claim
Sources
secondary
Hugging Face Daily Papers entry