Demystifying Agent Skills: Why They Work — Until They Don't
UC San Diego (Zhiyuan Jiang, Mengdi Wang, Yijiang Li et al.)
Controlled experiments across benchmarks, agent harnesses, and LLMs show skills beat Workflow Memory by 6.06 points but retrieval precision collapses from 29.6% to 3.3% as skill pools grow from 5 to 100. Procedural anchoring accounts for 65.7% of skill wins versus 4.5% from explicit knowledge injection.
Why it matters
Top paper on HF Daily Papers for Aug 19 with 149 upvotes — gives the first systematic ablation of when and why agent-skill libraries actually help, and surfaces the scaling cliff practitioners hit.
Importance: 3/5
HF Daily ≥100 upvotes