Yandex merges VLM and LLM into a single omnimodel powering Alice AI
Yandex
Yandex engineers detail how the separate text LLM and vision model behind Alice AI were unified into one omnimodel via a mixture-of-experts architecture, with sequential omni-pretraining plus SFT/RL alignment and ~300 tracked benchmarks. Side-by-side evals show it beating Yandex's own February model 57-43 and Qwen 3.5 397B (no-thinking) 59-41, while trailing Gemini 3.1 Pro and Qwen 3.5 thinking.
Why it matters
first detailed public account of the architecture behind Alice AI's flagship multimodal model
Importance: 3/5
flagship-model architecture disclosure from Yandex with competitive evals