Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
Makes the case that Evolution Strategies are a distinct post-training paradigm for LLM reasoning rather than a memory-efficient GRPO stand-in: ES achieves broader reasoning diversity (verifier-projected Jensen-Shannon diversity correlates with Pass@K), improves Pass@1 while exceeding GRPO on Pass@K without entropy collapse, and its gains come from a sparse subset of high-magnitude parameter updates, suggesting catastrophic forgetting is not a concern. Proposes a sequential GRPO-ES schedule combining both.
Why it matters
ES repositioned as a first-class RLVR alternative with better Pass@K coverage
Importance: 2/5
fresh arXiv drop, modest traction (15 upvotes)
Sources
official
arXiv 2608.27351
official
HF Papers — 15 upvotes, Aug 28