#performance
- Hugging Face Transformers: Async Continuous Batching Achieves 22% Inference Speedup Hugging Face tools
- Dharma-AI lifts GPU-cluster utilization 53.6% → 87.0% by encoding physical constraints into the scheduler Dharma-AI tools
- llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning llama.cpp (GGML) tools
- llama.cpp b10603-b10615 — GLM-4.5-Air MTP, Deepseek 4 -sm tensor, Metal flash-attn vec tuning for M1 Pro/M2 Ultra/M5 Max ggml-org tools