DataFlex-RL: evaluation platform finds RLVR data policies don't beat uniform sampling

Peking University

Research official + media 2 src. ~1 min

A controlled evaluation platform comparing 13 data-policy configurations for RLVR under a common GRPO recipe (Qwen2.5-7B-Base and Llama-3.1-8B-Base, 12 matched seeds, 12 benchmarks). No rollout-selection, reweighting, or adaptive domain-mixing method achieved a reproducible improvement over uniform GRPO, and benchmark-subset rescaling flipped rankings with a correlation of -0.33.

Why it matters

A careful negative result against a popular research direction: the fancy data-curation policies widely used in RLVR pipelines show no reproducible gain over plain uniform sampling when properly seeded and controlled.

Importance: 3/5

Careful negative result against a popular research direction; top HF Daily paper of the day

Sources