Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Research official 2 src. ~1 min

Makes the case that Evolution Strategies are a distinct post-training paradigm for LLM reasoning rather than a memory-efficient GRPO stand-in: ES achieves broader reasoning diversity (verifier-projected Jensen-Shannon diversity correlates with Pass@K), improves Pass@1 while exceeding GRPO on Pass@K without entropy collapse, and its gains come from a sparse subset of high-magnitude parameter updates, suggesting catastrophic forgetting is not a concern. Proposes a sequential GRPO-ES schedule combining both.

Why it matters

ES repositioned as a first-class RLVR alternative with better Pass@K coverage

Importance: 2/5

fresh arXiv drop, modest traction (15 upvotes)

Sources

official arXiv 2608.27351