#pretraining
- Video Generation Models are General-Purpose Vision Learners Google DeepMind research
- ByteDance Seed paper studies how high-quality domain data should repeat when scaling LLMs ByteDance research
- Yandex open-sources Alice AI Search Pretrain, the base model behind AI answers in Search Yandex models-llm
- NCP-ArchPreview: latent-space language models at 8.9B via Next Concept Prediction Shanghai AI Lab research
- Scalable Visual Pretraining for Language Intelligence research
- SMELT: looped MoE transformers match baseline scaling at matched compute ByteDance Seed research
- Understanding Reasoning from Pretraining to Post-Training research
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Meta AI research