ByteDance Seed paper studies how high-quality domain data should repeat when scaling LLMs
ByteDance
A ByteDance Seed team posted Scaling Domain Data Repetition in LLM Pretraining (arXiv 2608.14071, HF paper-page submission Aug 17, arXiv listing Aug 14). At fixed tokens-per-parameter the optimal repetition count mildly increases with model size; optimal repetition is strongly negatively correlated with validation loss across domains; counts tuned on smaller proxy models at the same TPP transfer to larger models.
Why it matters
Concrete, transferable rule for tuning how much to repeat scarce high-quality domain data as LLMs grow — directly actionable for training domain-specialised models on small corpora, since the recipe scales without retuning the search.
Importance: 3/5
open-weights / GA marker