DataFlex-RL: evaluation platform finds RLVR data policies don't beat uniform sampling
Peking University
A controlled evaluation platform comparing 13 data-policy configurations for RLVR under a common GRPO recipe (Qwen2.5-7B-Base and Llama-3.1-8B-Base, 12 matched seeds, 12 benchmarks). No rollout-selection, reweighting, or adaptive domain-mixing method achieved a reproducible improvement over uniform GRPO, and benchmark-subset rescaling flipped rankings with a correlation of -0.33.
Why it matters
A careful negative result against a popular research direction: the fancy data-curation policies widely used in RLVR pipelines show no reproducible gain over plain uniform sampling when properly seeded and controlled.
Importance: 3/5
Careful negative result against a popular research direction; top HF Daily paper of the day