Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

Research official 1 src. ~1 min

Formalizes a 'safety trilemma': for dual-use tasks, LLM safeguards cannot simultaneously offer useful capability, reliable safety, and open access, because an attacker can copy or imitate any prompt/context evidence a legitimate user would present. Derives a mathematical worst-case floor on attacker assistance and proposes 'trusted credentials', hard-to-copy, use-tied evidence, as the necessary complement to prompt-based safeguards.

Why it matters

A theoretical result with direct implications for how labs design refusal and access-control policies for dual-use capabilities, arguing today's copyable-context safeguards are structurally insufficient.

Importance: 2/5

Notable theoretical safety paper, single official arxiv source.

Sources