#rlvr
- VibeThinker-3B Reaches Frontier-Level Reasoning Benchmarks via Curriculum RL WeiboAI research
- Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs research
- TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Tencent research
- DataFlex-RL: evaluation platform finds RLVR data policies don't beat uniform sampling Peking University research
- DRPO: Rethinking Divergence Regularization in LLM Reinforcement Learning Tencent Hunyuan research
- Ring-Zero: Scaling Zero RL to 1 Trillion Parameters with Emergent Reasoning Behaviors Ant Group research
- TRL v1.11.0 rewrites trl vllm-serve (~10x smaller) and ships AsyncDistillationTrainer Hugging Face tools