EvoSafeHarness: Automatically Evolved Safety Harnesses for LLM Agents

Research official 1 src. ~1 min

A framework that synthesizes deployment-specific agent safety harnesses — jointly searching a natural-language policy and executable enforcement logic — for a frozen model in a target domain. On AgentDojo it reaches 82.8% utility at 0% attack success rate, roughly double CaMeL's utility at that operating point, and transfers unchanged to unseen suites.

Why it matters

Shows one-size-fits-all agent defenses are the bottleneck: per-model, per-domain harness search roughly doubles safe utility

Importance: 2/5

Notable agent-safety paper with strong AgentDojo results

Sources