NeoHorse-1: recursive self-improvement via agentic post-training
TokenRhythm
A family of agent-native models trained through agentic post-training: a heterogeneous model pool with intelligent routing records full harness-context interactions, converts them into training data, and closes an evaluation-selection-update loop toward harness-mediated recursive self-improvement. Routing signals organize a three-stage SFT curriculum plus routing-guided on-policy distillation; post-training lifts macro-average from 58.94 to 64.87 at 4B, narrowing the gap to the 9B base model.
Why it matters
367 upvotes on HF Daily Papers (Sept 9) — the most upvoted paper of the day. A concrete, open-weights mechanism for RSI where the system converts its own routed interaction traces into the next training round, with code and checkpoints released.
Importance: 4/5
Notable paper + 367 upvotes on HF Daily Papers; open-weights recursive self-improvement