-
llama.cpp v0.3.0 — first tagged stable release of the b10621 nightly; dots3-note multimodal, GLM-4.5-Air MTP, DeepSeek 4 tensor-split, ggml v0.22.0
ggml-org
tools
-
llama.cpp v0.4.0: lazy tensor reading, per-slot context limits, video input
tools
-
llama.cpp June 16 Builds: Eagle3 Speculative Decoding, Vulkan UMA Memory, NVFP4 Fixes
tools
-
llama.cpp b9754: Real-Time Model Load Progress via SSE and PEG Grammar Parser
tools
-
llama.cpp Builds b9830–b9837: DFlash v2, MiniCPM5 Parser, --reasoning-preserve Flag
ggml-org
tools
-
Ollama v0.31.2: MLX small-batch matmul kernel, llama.cpp build 9840, CUDA updates
tools
-
llama.cpp b9967–b9969: Adreno GPU Acceleration and OpenAI-Compatible Null Sampling
tools
-
llama.cpp Adds Tencent Hunyuan 3, Minimax2 Eagle3 Speculative Decoding, and SYCL Fused MoE
tools
-
llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning
llama.cpp (GGML)
tools
-
llama.cpp rolls up Aug 22 backend and model fixes
tools
-
llama.cpp b10603-b10615 — GLM-4.5-Air MTP, Deepseek 4 -sm tensor, Metal flash-attn vec tuning for M1 Pro/M2 Ultra/M5 Max
ggml-org
tools
-
llama.cpp b10660 adds Qwen3.8-Flash-Next (qwen4exp) architecture: hyper-connections, gated delta net layers, MoE, PLE n-gram embedding
llama.cpp
tools
-
llama.cpp nightly burst b10668-b10679 includes Vulkan wrong-token fix
llama.cpp
tools
-
llama.cpp brings MTP speculative decoding to recurrent Qwen models
ggml
tools
-
llama.cpp Sep 1-2 wave: +4.9% generation from n-gram lookup fix, fused CUDA MoE reduction
ggml
tools
-
llama.cpp v0.4.1 adds Maple 20B, Tencent Hy 4, and Spark2.5 support
ggml
tools
-
llama.cpp patches remotely exploitable use-after-free in llama-server RPC endpoint
ggml
tools