#safety
- Claude Fable 5 and Claude Mythos 5: Anthropic's Most Capable Model Goes Public Anthropic models-llm
- US Government Orders Anthropic to Disable Claude Fable 5 and Mythos 5 Globally Anthropic industry
- OpenAI Previews GPT-5.6 Family: Sol, Terra, and Luna in Government-Gated Limited Release OpenAI models-llm
- A Global Workspace in Language Models Anthropic research
- OpenAI unveils GPT-Red, an internal automated red-teaming model for prompt-injection defense OpenAI research
- OpenAI says pre-release models broke out of test sandbox and breached Hugging Face OpenAI research
- Anthropic CEO Clarifies: Company Does Not Oppose Open-Weight Models Anthropic industry
- Anthropic discloses three real-world incidents from Claude cybersecurity evaluations Anthropic research
- Exploration Hacking: LLMs Can Be Fine-Tuned to Strategically Resist RL Training research
- OpenAI Post-Mortem: How RLHF Reward Hacking Embedded Goblin Metaphors in GPT-5.x OpenAI research
- OpenAI Discloses Accidental Chain-of-Thought Grading in RL Training Across Six Models OpenAI research
- OpenAI Launches Daybreak: AI-Powered Vulnerability Detection Platform OpenAI tools
- US Congress Releases 269-Page 'Great American AI Act' Draft with 3-Year State Law Preemption industry
- Anthropic Staff to Meet White House Officials This Week to Negotiate Fable 5 Access Suspension Anthropic industry
- OpenAI Publishes Deployment Simulation: Predicting Model Behavior Before Release OpenAI research
- How Transparent is DiffusionGemma? Interpretability Study Closes the Gap to Autoregressive Models Google DeepMind research
- Anthropic Appoints Former Fed Chair Ben Bernanke to Long-Term Benefit Trust Anthropic industry
- Anthropic Research: Claude's Values Vary Significantly by Model Version and Language Anthropic research
- ElevenLabs Deploys Google DeepMind SynthID Audio Watermarking to Free-Tier Users ElevenLabs audio
- Google DeepMind and Isomorphic Labs detail joint bioresilience initiative Google DeepMind research
- Natural Language Autoencoders: Turning Claude's Thoughts into Text Anthropic research
- Anthropic Introduces Natural Language Autoencoders for Scalable LLM Interpretability Anthropic research
- Anthropic Eliminates Claude's Agentic Blackmail Behavior via 'Teaching Claude Why' Anthropic research
- GRAM: Modular Pretraining Makes Dual-Use Knowledge Physically Removable from AI Models AE Studio / Anthropic research
- Model Spec Midtraining: How Normative Self-Knowledge Improves Alignment Generalization Anthropic research
- SAE Interventions Are Unreliable: Suppressed Behaviors Recover Post-Intervention Hong Kong Polytechnic University research
- Google DeepMind Publishes AI Control Roadmap: Defense-in-Depth Against Misaligned Coding Agents Google DeepMind research
- GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents research
- SingGuard: Runtime Policy-Adaptive Multimodal LLM Guardrail with 56K-Example Benchmark inclusionAI research
- Anthropic Proposes Industry-Wide Cyber Jailbreak Severity Scale Anthropic research
- BadWAM: When World-Action Models Dream Right but Act Wrong research
- Anthropic's Project Pilot tests whether AI models can fly drones Anthropic research
- Meta Publishes Preparedness Report for Code World Model Before Open-Weight Release Meta research
- Google SynthID Reaches 100B+ Watermarked Assets; OpenAI and ElevenLabs Join C2PA Coalition Google DeepMind tools
- Cursor Launches Security Review Beta: PR Vulnerability Scanner and Scheduled CVE Agents Cursor tools
- Quantifying Faithful Confidence Expression in Large Reasoning Models Yale NLP research
- Anatomy of Post-Training: Using Interpretability to Audit and Fix Preference Data research
- Google DeepMind and Partners Launch $10M Multi-Agent AI Safety Research Fund Google DeepMind industry
- Anthropic Publishes First Public Record: 52,000-Person Survey on US AI Attitudes Anthropic research
- Claude Code v2.1.183: Auto Mode Safety Guards for Destructive Git and Infrastructure Commands Anthropic tools
- Institutional Red-Teaming: Deployment Rules, Not Just Model Weights, Causally Shape Multi-Agent Safety research
- Recursive Self-Improvement in AI: Survey of 1,250 Papers with Verification-Strength Taxonomy research
- Anthropic Launches Public 'Hard Questions' AI Accountability Initiative Anthropic industry
- Metacognition in LLMs: Foundations, Progress, and Opportunities — Yale Survey Yale NLP Lab research
- AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Stanford University research
- Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs research
- LLM Safety From Within (SIREN) University of Toronto CSSLab / McGill / LMU Munich research
- Yandex Extends ISO/IEC 42001 AI Safety Certification to Full Alice AI Model Family Yandex industry