Miles v0.1: a production-level RL post-training system

RadixArk

Research official 2 src. ~1 min

A full-stack open-source system for frontier RL post-training built on the slime design: SGLang rollout engines, Megatron-LM or PyTorch FSDP trainers, three weight-sync transports, plus LoRA RL, on-policy distillation, and diffusion-model support. Case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks on 64 GB300 GPUs, median step time 263 s.

Why it matters

2.7k GitHub stars within days; makes frontier-scale agentic RL reproducible outside frontier labs, with the first detailed public accounting of a 744B-scale agentic RL run.

Importance: 3/5

2.7k GitHub stars; first public 744B-scale agentic RL run details

Sources

official arXiv 2609.08368