COBRA-Skills: contextual bandit-guided evolution for agent skills

Chinese University of Hong Kong, Shenzhen

Research official + media 2 src. ~1 min

Treats LLM agent skill-library optimization as budgeted sequential optimization over an evolving candidate space, pairing contextual-bandit-guided prioritization with evidence-grounded skill evolution from execution feedback. Across six agent benchmarks and three models it matches or beats prior skill-optimization methods while cutting optimization cost by 55-58% relative to SkillOpt, using only 50 examples per benchmark.

Why it matters

Skill libraries are becoming the standard way to make agents improve over time; this shows the evolution loop can be made roughly half as expensive without losing performance.

Importance: 2/5

Halves the cost of agent skill-library optimization without losing performance

Sources