FlowBalance: verifier-grounded self-improvement for reasoning models

Research official 2 src. ~1 min

A self-improvement method that learns a normalized distribution over complete responses: a frozen training-time policy produces token-level log-probability gains, aggregated into a trajectory self-guidance score and calibrated against verifier-derived group advantage. It beats FlowRL on math reasoning with Qwen3-4B/8B, trains faster and more stably, and avoids the response-length collapse of direct on-policy self-distillation.

Why it matters

81 upvotes on HF Daily Papers; a concrete recipe for making the fragile on-policy self-improvement loop stable without dense human supervision.

Importance: 2/5

Solid methods paper, 81 upvotes on HF Daily Papers, below the 100-upvote bump bar

Sources