Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
National University of Singapore (Zhi Zheng, Wee Sun Lee et al.)
Replaces RL with evolution strategies for fine-tuning long-horizon LLM agents, needing only inference-level GPU memory and supporting trajectory-level credit assignment. On WebArena-Lite it improves a Qwen-3.5-27B agent by 6.69 points over a No-Skill baseline and wins 28 of 36 test-time automatic-heuristic-design settings.
Why it matters
96 HF upvotes (Aug 19) — offers a viable path for academic labs to fine-tune frontier-scale agentic policies on commodity GPUs, sidestepping the cost of agentic RL infrastructure.
Importance: 2/5
default
Sources
official
arXiv — Agentic ESOpt paper