NeoHorse-1: recursive self-improvement via agentic post-training

TokenRhythm

Research official 2 src. ~1 min

A family of agent-native models trained through agentic post-training: a heterogeneous model pool with intelligent routing records full harness-context interactions, converts them into training data, and closes an evaluation-selection-update loop toward harness-mediated recursive self-improvement. Routing signals organize a three-stage SFT curriculum plus routing-guided on-policy distillation; post-training lifts macro-average from 58.94 to 64.87 at 4B, narrowing the gap to the 9B base model.

Why it matters

367 upvotes on HF Daily Papers (Sept 9) — the most upvoted paper of the day. A concrete, open-weights mechanism for RSI where the system converts its own routed interaction traces into the next training round, with code and checkpoints released.

Importance: 4/5

Notable paper + 367 upvotes on HF Daily Papers; open-weights recursive self-improvement

Sources

official arXiv 2609.08183