#agentic-rl
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work Apodex research
- IBM releases Granite 4.2 reasoning LLMs (3B/8B/30B, Apache 2.0, 512K context) trained on ~15T tokens with GB200 NVL72 GRPO RL IBM models-llm
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work Apodex AI research
- OPID: On-Policy Skill Distillation Improves Long-Horizon Agent RL Institute of Automation, Chinese Academy of Sciences research
- EnvHarness: Awakening Static Worlds for Agent Learning Google research
- The Verification Horizon: No Single Reward Function Works for Coding Agents at Scale Qwen (Alibaba) research
- AgentOPSD: recursive self-distillation for agentic reinforcement learning research