←

Raschka: focus on post-training open-weight LLMs, not pre-training

Research official 1 src. ~1 min

Sebastian Raschka argues that teams without frontier-scale budgets get more value post-training an existing open-weight model than attempting their own pre-training, which costs far beyond even a few million dollars at 500B+ parameter scale. As a case study he cites Fireworks' Ember-1, post-trained from Kimi K3 for shorter reasoning and using roughly 40% fewer tokens at comparable quality in their evaluations. He notes token efficiency can be trained by putting token budgets into the reward objective, while cautioning that Fireworks has not disclosed enough of the recipe to confirm the method.

Why it matters

Practical evidence that post-training can buy large token-efficiency gains, shaping the economics of open-weight adoption.

Importance: 2/5

Practitioner analysis of post-training economics with a fresh case study

Sources