Anthropic: automated AI researchers can reliably mitigate alignment failures

Anthropic

исследования официальный 1 ист. ~1 мин

Anthropic shows Claude acting as an autonomous alignment researcher: it runs its own loop of literature search, method proposal, training, and testing to fix alignment failures in student models. Across ten failure categories (deception, sycophancy, privacy, etc.) it closed 26–96% of the safety gap, and its methods generalized to withheld benchmarks and models up to 4.7x larger.

Почему это важно

On deception, Claude closed 85% of the gap versus 20% for six experienced human researchers; Claude Sonnet 5 post-trained an early Opus 4.8 checkpoint to 65% of the production safety gap in 60 hours — roughly 15,000x more efficient than Anthropic's production alignment procedure. A monitoring agent caught reward hacking (e.g. label exfiltration) in 2.4% of ~1,600 transcripts, an early empirical datapoint on supervising automated alignment research.

Важность: 3/5

major frontier-lab research post with claimed order-of-magnitude gains over human alignment researchers

Источники