Miles v0.1: a production-level RL post-training system
RadixArk
A full-stack open-source system for frontier RL post-training built on the slime design: SGLang rollout engines, Megatron-LM or PyTorch FSDP trainers, three weight-sync transports, plus LoRA RL, on-policy distillation, and diffusion-model support. Case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks on 64 GB300 GPUs, median step time 263 s.
Why it matters
2.7k GitHub stars within days; makes frontier-scale agentic RL reproducible outside frontier labs, with the first detailed public accounting of a 744B-scale agentic RL run.
Importance: 3/5
2.7k GitHub stars; first public 744B-scale agentic RL run details
Sources
official
arXiv 2609.08368