llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning
llama.cpp (GGML)
Aug 18 work spans v0.1.2 pre-release (ggml sync to 0.20.2, MCP stdio docs + CORS defaults, integer tokenizer scores, refactor of Built-In Tools naming), b10488 (OpenVINO 2026.3), b10486 (LFM2 image tiling threshold fix), b10485 (ggml sync), and b10483 (cmake vendor:: alias targets). Aug 17 added b10472 (CUDA skip UMA override for HIP, #18159), b10470 (release.yml pushes tag explicitly), b10456 (SYCL q4_0→f32 20.21→158.19 GB/s on Arc 70), and b10455 (SYCL OPT_STEP_ADAMW/SGD).
Why it matters
The SYCL quantized-copy kernel fix delivers an ~8x throughput improvement for q4_0→f32 on Arc GPUs — a single kernel change that materially changes what local inference looks like on Intel discrete cards. OpenVINO 2026.3 and DGX Spark CUDA MMVQ tuning also widen the deployment matrix.
Importance: 2/5
default