#performance
- Hugging Face Transformers: Async Continuous Batching Achieves 22% Inference Speedup Hugging Face tools
- Dharma-AI lifts GPU-cluster utilization 53.6% → 87.0% by encoding physical constraints into the scheduler Dharma-AI tools
- llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning llama.cpp (GGML) tools
- llama.cpp b10603-b10615 — GLM-4.5-Air MTP, Deepseek 4 -sm tensor, Metal flash-attn vec tuning for M1 Pro/M2 Ultra/M5 Max ggml-org tools
- llama.cpp brings MTP speculative decoding to recurrent Qwen models ggml tools
- llama.cpp Sep 1-2 wave: +4.9% generation from n-gram lookup fix, fused CUDA MoE reduction ggml tools
- OpenClaw 2026.9.3 ships rehearsed updates and prompt-cache-preserving performance work OpenClaw tools