llama.cpp ReleasesOriginal · English

[Pre-release / continuous build] b11557

[Pre-release / continuous build] rpc : turn GGML_RPC_DEBUG into a verbosity level and add logs (#29544) * rpc : turn GGML_RPC_DEBUG into a verbosity level and add logs GGML_RPC_DEBUG is now parsed as a number: 0/unset…

Pre-release / continuous build: this is an experimental build, not a stable release.

rpc : turn GGML_RPC_DEBUG into a verbosity level and add logs (#29544) * rpc : turn GGML_RPC_DEBUG into a verbosity level and add logs GGML_RPC_DEBUG is now parsed as a number: 0/unset disables debug logs, 1-3 emit increasingly detailed output (events, per-command trace, transport detail). Non-numeric values fall back to 1. The duplicated env/macro blocks in ggml-rpc.cpp and transport.cpp are replaced by a shared log.h, which becomes the single choke point for all logging of the RPC backend: LOG_ERROR/LOG_WARN/LOG_INFO for unconditional severity logs and LOG_DBG/LOG_DBG2/LOG_DBG3 for the verbosity-gated ones. The transport files no longer need ggml-impl.h, and the server banner now also goes through the ggml logger (stderr). Missing logs are added on both the client (handshake, buffer ops, tensor transfers, graph computes, cache decisions) and the server (per-command dispatch, graph nodes), including the negotiated transport via the new socket_t::transport_name(). Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * rpc : align log levels with the documented verbosity semantics The first pass introduced GGML_RPC_DEBUG levels, but many call sites did not follow the documented semantics and some logs bypassed the macros entirely: - Route the remaining raw GGML_LOG_* call sites through the local LOG_* macros so that log.h stays the single choke point for RPC logging - Emit per-tensor transfers and per-command handler traces at level 2, and one-shot lifecycle events (backend and buffer type creation, buffer allocation, graph compute, graph cache eviction) at level 1 - Drop the duplicated enqueue trace in rpc_dispatcher::send, work() already logs every dispatched command together with its round-trip timing - Promote degraded-operation events to LOG_WARN: remote allocation failure, RDMA falling back to TCP, unexpected peer disconnect - Log graph cache eviction on both sides, since client and server must stay in sync for the incremental graph update to be valid - Distinguish Unknown command (opcode out of range) from Unhandled command (valid opcode with no handler) - Fix format specifiers (%zu for size_t, %u for device ids, 0x for hex pointers) and a missing newline in an error log - Document that llama.cpp applications drop ggml debug records unless the verbosity threshold is raised with -lv 5 Assisted-by: pi:llama.cpp/Qwen3.8-Flash-Next

Website: - https://llama.app

Attestations: - https://github.com/ggml-org/llama.cpp/attestations/54812036

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries - Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLLs - Windows x64 (CUDA 13) - CUDA 13.4 DLLs - Windows arm64 (CUDA 13) - CUDA 13.4 DLLs - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0)

openEuler: - DISABLED - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph)

UI: - UI

Original source

llama.cpp Releases

Content notes

Original publication and rights belong to the source.