←

Your Transformer Can Hold Two Thoughts at Once: evidence of linear superposition in LLMs

Research official + media 2 src. ~1 min

Proposes the Superposition Linearity Hypothesis: mixing two text streams as inputs yields an output distribution approximately equal to the blend of each stream's separate predictions. The trait diminishes during pretraining but is recoverable with light fine-tuning, and the authors demonstrate dual-stream guided decoding from a single forward pass.

Why it matters

Second-most-upvoted HF paper for Sep 25 (55 upvotes); a testable empirical account of superposition directly in activations rather than in SAE features.

Importance: 3/5

Highly-upvoted HF Daily paper with a new testable interpretability claim

Sources