llama.cpp ReleasesOriginal · English

[Pre-release / continuous build] b11528

[Pre-release / continuous build] meta : handle host views (#30217) * meta : handle views of tensors allocated on the host A view shares the memory of its view_src, so ggml-alloc never allocates a view in the buffer of…

Pre-release / continuous build: this is an experimental build, not a stable release.

meta : handle host views (#30217) * meta : handle views of tensors allocated on the host A view shares the memory of its view_src, so ggml-alloc never allocates a view in the buffer of the split it lands in - the scheduler copies the source into the split and the ops that use the view read that copy. The view node itself is a noop and does not need a split of its own, but the meta backend asserted when one was left inside a meta split: - ggml_backend_meta_get_split_state() dereferenced tensor->buffer->context - the graph rebuild mapped every node with ggml_backend_meta_buffer_simple_tensor() Accept such nodes when they are views of host tensors, which also generalizes the previous s_copy_main workaround. This fixes the assert hit by KV cache views when using --split-mode tensor with partial offload. Assisted-by: pi:llama.cpp/Qwen3.8-Flash-Next * archs : re-enable sm tensor for K2 Horizon * cont : add TODO and reference

Website: - https://llama.app

Attestations: - https://github.com/ggml-org/llama.cpp/attestations/54357448

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries - Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLLs - Windows x64 (CUDA 13) - CUDA 13.4 DLLs - Windows arm64 (CUDA 13) - CUDA 13.4 DLLs - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0)

openEuler: - DISABLED - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph)

UI: - UI

Original source

llama.cpp Releases

Content notes

Original publication and rights belong to the source.