#mech-interp
- A Global Workspace in Language Models Anthropic research
- How Transparent is DiffusionGemma? Interpretability Study Closes the Gap to Autoregressive Models Google DeepMind research
- The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models research
- Judge Circuits: Mechanistic Explanation of LLM-as-Judge Format Inconsistency research
- SMELT: looped MoE transformers match baseline scaling at matched compute ByteDance Seed research
- Anatomy of Post-Training: Using Interpretability to Audit and Fix Preference Data research
- GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning research
- Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs NAVER AI Lab research