#mechanistic-interpretability
- A Global Workspace in Language Models Anthropic research
- Natural Language Autoencoders: Turning Claude's Thoughts into Text Anthropic research
- Anthropic Introduces Natural Language Autoencoders for Scalable LLM Interpretability Anthropic research
- Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and a Fix research