llama.cpp ReleasesOriginal · English

[Pre-release / continuous build] b11408

[Pre-release / continuous build] ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479) * ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 On the SpacemiT X60, IME matrix acceleration only covered…

Pre-release / continuous build: this is an experimental build, not a stable release.

ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479) * ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 On the SpacemiT X60, IME matrix acceleration only covered Q4_0/Q4_1/Q4_K. Q8_0 had no IME1 kernel, and since the SpacemiT build sets GGML_CPU_REPACK=OFF there was no repack path compiled in either, so Q8_0 had no accelerated path at all and ran roughly ten times slower than Q4_0 for prefill on the same board. - add make_block_q8_0x16 and the Q8_0 repack entry: interleave the weights into the 16-column layout the IME1 vmadot sequence expects - add ime1::gemm_kernel_i8i8, an int8 x int8 IME1 kernel with a single-row and a 4-row A path; the 4-row path loads each B panel once and reuses it across 4 rows of A - add quantize_a_4row_i8 for the 4-row activation quantization - wire both into forward_mul_mat and the repack factory for Q8_0 - docs: mark Q8_0 as supported on X60 Correctness was checked against a quant-exact integer reference for K = 32 up to 4096, with a max relative error of about 1e-6, and by checking that generation stays coherent across several prompts. Tested on Milk-V Jupiter (SpacemiT X60), Bianbu 2.1.1, gcc 14.2, with Qwen2.5-0.5B-Instruct Q8_0. llama-bench -t 4 under taskset -c 0-3, 5 repetitions on an idle board: pp128 goes from 10.70 to 93.87 t/s. Q4_0 is unchanged at 106.40 -> 107.51 t/s, as expected since this does not touch that path. * ggml-cpu : move q8_0_16x32 decl to IME1 section * ggml-cpu : align q8_0 IME1 kernel assignments

Website: - https://llama.app

Attestations: - https://github.com/ggml-org/llama.cpp/attestations/52757256

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries - Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLLs - Windows x64 (CUDA 13) - CUDA 13.4 DLLs - Windows arm64 (CUDA 13) - CUDA 13.4 DLLs - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0)

openEuler: - DISABLED - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph)

UI: - UI

Original source

llama.cpp Releases

Content notes

Original publication and rights belong to the source.