#llama-cpp
- llama.cpp June 16 Builds: Eagle3 Speculative Decoding, Vulkan UMA Memory, NVFP4 Fixes tools
- llama.cpp b9754: Real-Time Model Load Progress via SSE and PEG Grammar Parser tools
- llama.cpp Builds b9830–b9837: DFlash v2, MiniCPM5 Parser, --reasoning-preserve Flag ggml-org tools
- Ollama v0.31.2: MLX small-batch matmul kernel, llama.cpp build 9840, CUDA updates tools
- llama.cpp b9967–b9969: Adreno GPU Acceleration and OpenAI-Compatible Null Sampling tools
- llama.cpp Adds Tencent Hunyuan 3, Minimax2 Eagle3 Speculative Decoding, and SYCL Fused MoE tools