-
Google DeepMind Releases Gemma 4 QAT Checkpoints: Sub-1 GB On-Device E2B Model
Google DeepMind
models-llm
-
Ollama v0.33.2 (pre-release) restores system dark mode and keeps the proxy live across catalog changes
Ollama
tools
-
Ollama v0.24.0: Codex App Integration and MLX Sampler Improvements
Ollama
tools
-
llama.cpp b9161/b9169: Codex CLI Compatibility and Qwen3A Multimodal Support
ggml-org
tools
-
Ollama v0.30.7: Hermes Desktop Support, Gemma 4 QAT, and Nemotron-3-Ultra
Ollama
tools
-
llama.cpp b9589–b9592: CUDA SSM Sync Fix and Mamba Memory Optimization
tools
-
Ollama v0.30.9: Cohere2Moe Support, Coding Agent Single-Token Output Bug Fixed
tools
-
llama.cpp June 16 Builds: Eagle3 Speculative Decoding, Vulkan UMA Memory, NVFP4 Fixes
tools
-
Ollama v0.31.1: Gemma 4 Nearly 90% Faster on Apple Silicon via MTP
Ollama
tools
-
Ollama v0.31.2: MLX small-batch matmul kernel, llama.cpp build 9840, CUDA updates
tools
-
llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning
llama.cpp (GGML)
tools
-
Ollama v0.33.1 adds Qwen3.8 Flash Next support and structured-output mlxrunner
Ollama
tools
-
llama.cpp v0.4.1 adds Maple 20B, Tencent Hy 4, and Spark2.5 support
ggml
tools
-
llama.cpp patches remotely exploitable use-after-free in llama-server RPC endpoint
ggml
tools
-
llama.cpp v0.5.0 stable: CUDA/Metal kernel work, video_url in server
tools
-
Ollama v0.34.4 pre-release fixes model-not-found flake, thinking-model structured outputs
tools