#local-inference
- Google Releases DiffusionGemma: 26B Open Model with 4× Faster Text Generation Google DeepMind models-llm
- PrismML ships Bonsai 2 27B: ternary weights near-lossless to the teacher PrismML models-llm
- CUDA-for-AMD-on-Windows stack runs CUDA apps on Radeon via ZLUDA tools
- llama.cpp b9754: Real-Time Model Load Progress via SSE and PEG Grammar Parser tools
- llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning llama.cpp (GGML) tools
- llama.cpp b10660 adds Qwen3.8-Flash-Next (qwen4exp) architecture: hyper-connections, gated delta net layers, MoE, PLE n-gram embedding llama.cpp tools
- Ollama v0.34.0-rc1: local models usable directly in ChatGPT Desktop tools