llama.cpp ReleasesOriginal · English

[Pre-release / continuous build] b11513

[Pre-release / continuous build] CUDA: improve top-k algorithm selection (#28713) * CUDA: radix top-k for large row counts Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select, gated on…

Pre-release / continuous build: this is an experimental build, not a stable release.

CUDA: improve top-k algorithm selection (#28713) * CUDA: radix top-k for large row counts Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select, gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts top-k from 1,671,253 launches / 5,761.8 ms to 2,329 / 941.8 ms. * CUDA: select the TOP_K implementation by shape Replace the nrows/ncols special case with the decision boundary from #28547 (as implemented in #29278): bitonic for short rows, radix select for several long rows, and DeviceTopK or CUB argsort for a single long row. The thresholds stay overridable at build time. Two refinements on top of that boundary: - bitonic stays in use for rows up to a padded 1024 while the rows fit in one wave of blocks (nrows <= number of SMs); radix select pays a fixed cost of about a dozen launches that only amortizes over more rows - with DeviceTopK available, it handles up to two rows Radix select now processes rows in chunks so its scratch memory stays bounded, and the bitonic path keeps its chunking. HIP and MUSA keep their previous thresholds. Add perf cases around the bitonic/radix crossover to test-backend-ops. * CUDA: make top-k comments less verbose * CUDA: remove the TOP_K width limit from supports_op * CUDA: use DeviceTopK for single-row TOP_K if available * CUDA: avoid ncols overflow in the TOP_K bitonic check * CUDA: share the row chunking helper between argsort and top-k * CUDA: do the TOP_K radix blocks_per_row math in int64_t * CUDA: rename GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK to GGML_CUDA_TOP_K_NROWS_THRESHOLD * CUDA: share one sort helper between the bitonic and CUB TOP_K paths * CUDA: update the TOP_K TODO, threshold and chunking comments * tests: add TOP_K cases that span several row chunks * CUDA: use int64_t col in the TOP_K radix loops, fix threshold comment * CUDA: limit TOP_K and ARGSORT support to ne[0] <= INT_MAX --------- Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com> Co-authored-by: Pranesh Gonegandla

Website: - https://llama.app

Attestations: - https://github.com/ggml-org/llama.cpp/attestations/54066267

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries - Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide

Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLLs - Windows x64 (CUDA 13) - CUDA 13.4 DLLs - Windows arm64 (CUDA 13) - CUDA 13.4 DLLs - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0)

openEuler: - DISABLED - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph)

UI: - UI

Original source

llama.cpp Releases

Content notes

Original publication and rights belong to the source.