Ollama v0.33.1 adds Qwen3.8 Flash Next support and structured-output mlxrunner
Ollama
v0.33.1 (Aug 26) adds MLX support for Qwen3.8 Flash Next, makes cmake compatibility patches idempotent, updates the bundled MLX and llama.cpp trees, and adds structured-output support to the MLX runner (including avoiding Metal GPU timeouts when loading models from slow storage).
Why it matters
Qwen3.8 Flash Next on Apple Silicon closes the gap between the desktop Ollama build and the latest Qwen lineage; structured outputs in mlxrunner unblock constrained-decoding use cases on local Macs.
Importance: 2/5
default
Sources
official
Ollama v0.33.1 release