FlowBalance: verifier-grounded self-improvement for reasoning models
A self-improvement method that learns a normalized distribution over complete responses: a frozen training-time policy produces token-level log-probability gains, aggregated into a trajectory self-guidance score and calibrated against verifier-derived group advantage. It beats FlowRL on math reasoning with Qwen3-4B/8B, trains faster and more stably, and avoids the response-length collapse of direct on-policy self-distillation.
Why it matters
81 upvotes on HF Daily Papers; a concrete recipe for making the fragile on-policy self-improvement loop stable without dense human supervision.
Importance: 2/5
Solid methods paper, 81 upvotes on HF Daily Papers, below the 100-upvote bump bar