llama.cpp ReleasesOriginal · English

[Pre-release / continuous build] b11516

[Pre-release / continuous build] hexagon: fix IM2COL patch-embed DMA ring overflow (#30189) * hexagon: fix IM2COL patch-embed DMA ring overflow The exact-tiling (stride == kernel, no pad/dilation) IM2COL DMA kernel…

Pre-release / continuous build: this is an experimental build, not a stable release.

hexagon: fix IM2COL patch-embed DMA ring overflow (#30189) * hexagon: fix IM2COL patch-embed DMA ring overflow The exact-tiling (stride == kernel, no pad/dilation) IM2COL DMA kernel issues IC*KH DDR->VTCM descriptors per output row without checking the return value of dma_queue_push(), and then pops IC*KH times. The per-thread DMA ring holds 256 entries and a push into a full ring returns false and drops the transfer, so for IC*KH > 255 the remaining rows of the VTCM staging buffer were never written and stale data (often NaN/inf) leaked into the output. Solution is to retire the oldest descriptor when the ring is full, just as the blocked kernel in the same file already does, and wait with dma_queue_flush(). For testing, added exact-tiling test cases with IC*KH > 256 (2D 1x1, 2D 2x2 patch embed and 1D, F16 and F32 dst), which fail on HTP without this fix. * Apply suggestion from @max-krasnyansky --------- Co-authored-by: Max Krasnyansky

Website: - https://llama.app

Attestations: - https://github.com/ggml-org/llama.cpp/attestations/54222989

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries - Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLLs - Windows x64 (CUDA 13) - CUDA 13.4 DLLs - Windows arm64 (CUDA 13) - CUDA 13.4 DLLs - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0)

openEuler: - DISABLED - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph)

UI: - UI

Original source

llama.cpp Releases

Content notes

Original publication and rights belong to the source.