llama.cpp v0.5.0 stable: CUDA/Metal kernel work, video_url in server

Tools official 1 src. ~1 min

llama.cpp tagged v0.5.0 (Sep 23) with CUDA conv2d implicit-GEMM acceleration, Metal MoE and SSM_CONV fusion, multi-address server binding, draft-model support for Gemma4 DSpark, new models (MiMo-V2.6, Nemotron MTP/H, Qwen4Exp, HRM-Text/DFM Mimir 1B), and a ggml bump to 0.25.0. The server also gained the OpenAI-standard video_url content type and data:video/* URIs.

Importance: 2/5

Major-version stable release of a core local-inference engine

Sources