#monitorability
- OpenAI Discloses Accidental Chain-of-Thought Grading in RL Training Across Six Models OpenAI research
- How Transparent is DiffusionGemma? Interpretability Study Closes the Gap to Autoregressive Models Google DeepMind research
- Anthropic Launches Public 'Hard Questions' AI Accountability Initiative Anthropic industry