Pre-release / continuous build: this is an experimental build, not a stable release.
llama : fix unexpected graph reallocation in the k-pool models (#29958) * llama : fix unexpected graph reallocation in the k-pool models Both k-pool models built a graph shape that depends on state the full-context reserve cannot know: - qwen4exp branched on inp->cache_safe, which turns false as soon as llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA layers swapped scatter+gather for fill+concat and dropped the new_pool_rep leaf, so the decode graph had 12 fewer nodes than the reserved one - glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the TG decode built the gather shape (7564 nodes) while the last reserve, the PP one, had the dense shape (7762 nodes) Either mismatch forces a decode-time re-reserve that drops the worst-case sizing and bakes in the current state, so the next state growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.: GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench \ -hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2 \ -c 32768 -pps -kvu Always scatter+gather the pooled keys, and pick gather from context constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds n_sel. Every graph of a context then shares one shape, which the reserve covers, and the dense path measured faster than the gather path at 2.5k and 16k context. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * llama : drop the unused k-pool cache_safe graph API The k-pool graphs no longer branch on cache_safe, so nothing reads get_kpool_cache_safe() or the conditional new_pool_rep any more: both models always pass the scatter target, which set_input_kpool now requires instead of merely preferring. Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only helper. The layout and state flag itself stays, it still decides which pools a layout with shared cells must re-pool. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * tests : add a shared-seq graph reserve regression test Decode a prompt into seq 0, share its cells with seq 1 via llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep decoding both sequences. For the k-pool models sharing clears cache_safe, which changes the graph topology while the pools keep growing, so a scheduler that re-reserves with the current state instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1. The test registration sets that flag, and the test aborts on both k-pool models before 2220411ec1. kimi-linear and minimax-01 are skipped: they reserve the final pp graph with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so every multi-seq graph has a different layout and re-reserves by design. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * cont : add TODOs * cont : fix comment * cuda: match the moe weighted reduction on empty ubatches ggml_cuda_match_moe_weighted_reduction rejected tensors with zero rows. A ubatch without outputs shrinks the last layer to zero rows through inp_out_ids, so graph_optimize dropped its alloc dep there and the scheduler graph lost one node compared to the reserved one. The scheduler then re-reserved at the size of that ubatch, and the next ubatch with the same node count but larger tensors aborted under GGML_SCHED_DEBUG_REALLOC=1. The compute loop already skips empty nodes before trying any fusion, so the guard only made the alloc deps depend on the row count. * tests: build the rollback test only where internal symbols link The shared-seq case calls llm_arch_from_string, which libllama does not export through LLAMA_API, so linking test-recurrent-state-rollback fails on Windows with shared libraries. Its build now sits in the NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and the test registration it already lives under. * tests: skip archs by name in the shared-seq reserve test The skip of kimi-linear and minimax-01 went through llm_arch_from_string, which libllama does not export through LLAMA_API, so the test could not link on Windows with shared libraries. It now compares the general.architecture string directly, and the test builds on every platform again. --------- Co-authored-by: PascalWebsite: - https://llama.app
Attestations: - https://github.com/ggml-org/llama.cpp/attestations/52789660
macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED - macOS Intel (x64) - iOS XCFramework
Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries - Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLLs - Windows x64 (CUDA 13) - CUDA 13.4 DLLs - Windows arm64 (CUDA 13) - CUDA 13.4 DLLs - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0)
openEuler: - DISABLED - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph)
UI: - UI