Ollama v0.33.1 adds Qwen3.8 Flash Next support and structured-output mlxrunner

Ollama

Tools official 1 src. ~1 min

v0.33.1 (Aug 26) adds MLX support for Qwen3.8 Flash Next, makes cmake compatibility patches idempotent, updates the bundled MLX and llama.cpp trees, and adds structured-output support to the MLX runner (including avoiding Metal GPU timeouts when loading models from slow storage).

Why it matters

Qwen3.8 Flash Next on Apple Silicon closes the gap between the desktop Ollama build and the latest Qwen lineage; structured outputs in mlxrunner unblock constrained-decoding use cases on local Macs.

Importance: 2/5

default

Sources