llama.cpp ReleasesOriginal · English

v0.6.0

## Overview llama.cpp v0.6.0 introduces the new `llama_batch_ext` extended batch API (with `llama_process`) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash…

Overview

llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.

Highlights

  • New llama_batch_ext extended batch API with llama_process(), supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669
  • New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model #27773, and the Clef decision model, fully supported with both text and vision #29831 #29969
  • Qwen4Exp: high-quality support is now available, with MTP speculative decoding (~1.5x decode speedup on DGX Spark) and various correctness fixes #29761 #29751
  • llama and server: new /v1/systemone API supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818
  • Metal: new tensor API flash attention kernel for F16 KV #29570
  • Metal: new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs #29869
  • New llama_prefetch_rows() using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599

API changes

  • include/llama.h: new llama_batch_ext batch API with llama_embd, llama_process() and llama_process_type #24669, new llama_get_causal_attn() #28876, session formats bumped to LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4
  • include/llama-cpp.h: added llama_batch_ext_ptr and deleter for the new extended batch API #24669
  • tools/mtmd/mtmd.h: mtmd_get_memory_usage() now returns an mtmd_memory_usage struct with image_max_tokens and use_non_causal #29773
  • tools/server: new /v1/systemone endpoint for decision models #29818 and /v1/embeddings now accepts typed vision/audio/video content #29556

New models

  • GLM-5.3-Flash (GLM5-Next): 320B KDA/DSA hybrid text+vision model with mHC and MoE #27773
  • Clef decision model, fully supported with both text and vision #29831 #29969
  • Ling 3.0 VL, folded into the BailingMoeV3 architecture #29151
  • Nimble decision model #29844
  • Registered Lfm2BidirectionalForMaskedLM for LFM2.5-Encoder-230M/350M #29862
  • Added classifier_pooling support for rerankers #29627

Core changes

  • Migrated examples, speculative decoding, mtmd and server to the new llama_batch_ext API #29385 #29601; batches now accept both embd and raw tokens #29622
  • Added llama_prec_policy and a model-driven W4A4 (NVFP4/MXFP4) mul_mat path #24364
  • Qwen4Exp: halved indexer score memory #29825, optimized mask constructions #29824, re-enabled the -sm tensor #28569; GLM5-Next: unique scatter rows for dead indexer slots #29745
  • KV cache: fixed restoring mismatched KV cache rotation #28498, fixed K/V and recurrent state cleanup after failed restores #27530 and an invalid assert in recurrent memory #29799
  • Speculative decoding: probabilistic sampling for simple draft and MTP #27694, fixed n-gram drafts rejected at temp > 0 after truncation #29924, preserved original batch order for layer inputs #29019, stop accepting draft tokens at EOG #29638
  • k-pool models: fixed unexpected graph reallocation #29958 and clamped kpool re-pool bound to existing pools #29805
  • Fixed tensor split for fused qkv with uneven K/V head sizes #29294, gather recurrent states once so the reserve covers every split #29856, and properly handle KV on training #28520
  • Context: do not re-reserve the scheduler when toggling causal_attn #28751
  • DFlash drafts: write Gemma embedding scale during conversion #29802 and add dflash support for MiMo #29650

Multi-modality changes

  • Cap max_image to n_ubatch for non-causal models #29773
  • Fixed the mel preprocessor in LFM2 audio #29403
  • Migrated input processing to llama_batch_ext #29385

Server changes

  • New /v1/systemone decision-model API with dedicated decision pipeline #29818, extended to the nimble decision model #29844
  • Support vision input for Clef #29969
  • GET /v1/models and GET /models now report model input/output modalities in a new architecture object #29987
  • /v1/embeddings: accept typed content (vision/audio/video) input #29556 and return HTTP 400 for invalid embedding requests #29060
  • Allow RANK pooling batch splitting for causal LLM rerankers (Qwen3, Qwen3-VL) #28876
  • Reject partial media truncation #24076
  • Logging: self-contained colors and split child commands from logs in router mode #29895, allow preset to set log file #29334, fixed dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list #29938
  • Fixed laya abort by limiting n_batch to n_ubatch #29903
  • Remove the built-in UI's service worker when the UI is not served #29565

UI changes

  • New model download pipeline #27959, Hugging Face Hub data layer #27947, model memory-fit estimation #27957 and model id grammar for sidecars, quants and capability parsing #27946
  • Type-safe API types, fetch helpers and download-ready models store plumbing #29582
  • Shared model display primitives #29644
  • Fixed missing svg use and animation elements in preview and download #28962
  • Use toLocaleString() formatting consistently across chat message statistics #27990

ggml changes

  • ggml updated to v0.26.0: a new alloc_buffer_n/get_alloc_size_n buffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.

Assets

Nightly build: b11429

More info

Changelog since v0.5.0

d81235049 llama.cpp : bump version to 0.6.0 (#29997) 4d60b4d08 common, server : report model input/output modalities in GET /models (#29987) c06f84160 sync : ggml f05c8b278 ggml : bump version to 0.26.0 (ggml/1652) e117148a4 CUDA: make the alloc_deps check batch independent (#29986) 6c59c4007 vulkan: fix Flash Attention shmem write out of bounds (#29988) 3c9e747f7 vulkan: revert mul_mat_id tile selection PR #29182 (#29936) b809b886d cuda: use the vector lightning indexer kernel on MUSA (#29990) 994e8f222 ci : add "Require Docker" flag to make-release workflow (#29989) 9d853bb36 webui: Use toLocaleString() format consistently across chat message statistics (#27990) 8f9ae20c8 ci : disable failing test on virtual Metal device (#29993) 9871df591 server: support vision input for Clef (#29969) 8b2fbaf32 CUDA: Optimize accumulation in mmq for NVFP4 type (#29857) 2ed93db47 ci : disable unused qemu in docker build (#29984) 8e1642198 server: reject partial media truncation (#24076) 806eee984 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (#29591) b3daa077a vulkan: sparse flash attention for quantized K/V (#29639) c173a53bd llama : fix unexpected graph reallocation in the k-pool models (#29958) 210791069 kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata (#28498) e5983d670 ci : winget urls must be separate strings (#29978) 4ca6b76f0 ci : fix docker workflow permissions (#29979) 9f12cd4a4 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479) ebe18bee5 vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (#29912) 8216c8462 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (#29483) 1b43d3116 cuda: tile the lightning indexer over keys and tokens for 4 heads (#29901) a3a1c4747 metal : few-row MMA mat-mul (#29869) 9d3aba6b5 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (#29633) d89651a7b CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (#29435) a7fb71fab log, server: self contained colors, split child commands from logs in router mode (#29895) 0bb496dbd llama: support both embd + raw tokens in batch (#29622) 2ca15f540 CUDA: refactor swizzling code (#29612) a7b94df2c ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806) 0eb6d9a81 cuda : move neu_padded to where it is used (#29940) 2e7c58c54 ci : windows llvm build requires ninja multi-config (#29959) 7f2dd88b0 ci : add windows arm64 vulkan release (#29954) bf79dbbcd AGENTS.md : revamp (#29656) dbe4c3ed4 chat-peg-parser : clear current_tool when pending_tool_call is reset (#29942) 46847e615 ci : set default permissions (#29945) 2bc563573 cuda : move blocks_per_col to where it is used (#29939) dd266785c CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941) 16c163d56 vulkan: fix rdna4 mat_vec tuning (#29934) 050439614 imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891) 8330e9696 spec : fix n-gram drafts rejected at temp > 0 after truncation (#29924) 6716df694 common : prepare load_from_models_dir() for path conversion (#29674) bf9a0ccce server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (#29938) 0faee5004 ci : pushing tag needs deploy key (#29937) f98b31c67 ci : improve release flow (#29913) 11fe02151 webgpu: add f16 support to fill/set_rows (#29897) 836d57176 mtmd : fix deprecated strdup warning on Windows (#29863) eec18f5d3 vendor : update cpp-httplib to 0.59.0 (#29886) 1537a0a8b server : fix laya abort by limiting n_batch to n_ubatch (#29903) edd6e2bbd common : add common_is_tty() helper and fix deprecated warnings on Windows (#29860) 9bf55f4a3 chat : honor json_schema in Ling 3.0 parser (#29813) a55e952b8 ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (#29904) 436f6f89e graph: gather the recurrent states once so the reserve covers every split (#29856) b92761a51 ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852) cb7934c52 model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (#29862) 889edf43d qwen4exp : halve the indexer score memory (#29825) 99b95488c model: add support for clef decision model (text-only) (#29831) bed0a8566 CUDA: fuse shared experts into MMVQ (#29184) 4ebdf2c74 ci : use t4-medium for cuda jobs (#29842) 1fb7ef3e3 spec : add probabilistic sampling for simple draft and MTP (#27694) 134b2bb75 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (#27663) 2923cf286 ggml-quants : avoid invalid rounding in qkx3 scale search (#29817) dd4c286f3 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096) 46ca246de model: support nimble decision model (#29844) d8fbd2583 readme : add cmd install commands (#29850) 926862e57 metal : add tensor API flash attention kernel for F16 KV (#29570) a4cb4c61f llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (#29818) 70849ee82 common : remove fs_open_ifstream() by using u8path() (#29841) 8d81559fa llama : silence unused-result warnings (#29839) 6805ae35d llama : use GGML_ABORT instead of throw (#29840) a8c9a4e7c opencl: use sigmoid f16 for bf16 (#29787) 392ded654 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (#29186) 9e258a6e0 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (#28531) b93328954 sycl: large register file for D=512 FA vec kernels (#29062) c328acc91 sycl : do not use slow oneDNN reference matmul and fattn (#28985) 4e2713c16 qwen4exp : optimize mask constructions (#29824) 631109b34 ggml : add alloc_buffer_n to buffer type interface (#23671) 254b17730 ci : fix missing zdnn backend check (#29837) fb4b2737a vulkan: add logging to pipeline compile issues (#29794) 207bdab95 pyproject : add linux platform marker to uv torch source (#29177) 5fc4f3c8c hexagon: install rebuilt HTP skels (#29828) 159c651f5 qwen4exp: fix tests (#29819) a868c3e3c hexagon: add q2_k and q3_k quant type support (#29717) ec7630a64 CUDA: fix 2 broken Volta FA cases (#29803) 78e2964c2 llama: refer to segment documentation [no ci] (#29074) f1cee9941 common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (#29816) 68e79bd8c skill: note about model-specific CLI arguments + testings (#29808) e358d5917 ci: fix Fusion / metal by updating the qwen4exp baseline (#29812) 81e39ad34 llama : clamp kpool re-pool bound to existing pools (#29805) dcd387a41 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685) d775ebf36 server: return HTTP 400 for invalid embedding requests (#29060) 2b36825cb convert : write Gemma embedding scale for DFlash drafts (#29802) 42d958167 cuda : route sm70 to the Turing MMVQ nwarps table (#29753) 13b4d7135 metal : release temporary private transfer buffers (#29777) 4b1622afb webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (#29358) 869034b4b llama : fix invalid assert in recurrent memory (#29799) b56f34ab1 CUDA: Handle compute type for NVFP4 on cublass path (#29173) c061df198 Qwen4Exp: add MTP (#29761) 66e0c17ee llama: fix qwen4exp (#29751) 767767850 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792) 552f18f91 mtmd: cap max_image to ubatch for non_causal models (#29773) 5503b04b0 meta: clear inactive AllReduce shards with FILL, not SCALE (#29793) def4d406a jinja : skip copying loop scope unless a loop filter needs it (#29776) 32dd62ee6 llama-mmap : avoid a second full-size copy of each tensor with direct-io (#29749) f11d642a2 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (#29572) 3aa0ce9bc hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (#29785) b0aca3c65 BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640) b8f96c3e8 common : add LLM-jp-4.1 Harmony dialect handler (#29681) 3ec4df42d opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (#29698) db33d3cb8 vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734) 7dad6db85 llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229) 2232bc8b5 metal : use bf16 math for mxfp4 mul-mat (#29770) 79625e056 llama-bench : fix docs (#29464) 66bcc2770 docs : refresh CPU ops support matrix (#29666) 10f340d1a model : re-enable -sm tensor for qwen4exp (#28569) 0c1e57098 webgpu: fix SSM_SCAN binding aliasing (#29750) f7b384c1e ggml-opencl : replace alloca() with std::vector (#29765) f872b5911 cuda: guard the iq4_nl dequantize row kernel against short rows (#29683) a4d880fd5 Hexagon: optimize ALLREDUCE with support for safe scatter mode (#29757) feb9a3d6d args: fix cli download mmproj arg (#28977) 4453b535f llama : preserve original batch order for speculative decoding layer inputs (#29019) 4f31296a9 test-llama-archs : toggle causal_attn to catch graph shape changes (#29724) b016f461b convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324) 81ff93ea1 llama: properly handle KV on training (#28520) 60e9cf7a7 batch: migrate the rest of examples to llama_batch_ext (#29601) 05af0d2b1 glm5-next: give dead indexer slots unique scatter rows (#29745) 2149c00f4 ggml/gguf : fix integer overflow (#29384) 876c75b1f codeowners : remove former ZenDNN owner (#29747) b04642061 cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (#29722) 22bdcc4cd mimo : support dflash (convert + feature extraction) (#29650) ca2e2037b jinja : support coerced array attributes (#29574) bdeb855b3 ggml-et : remove useless alloca() (#29663) 3b3d022b8 ci : fix Models Backend Check by shortening the hrm_text fixture (#29744) 185103dcf llama: llama_prefetch_rows (#29599) 2090f60f0 ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (#29675) 90c908d06 cpu: accept BF16 in src1 of mul_mat (#28937) 8df332de1 model-conversion : add --add-bos to run org model script (#29558) 4a096b8ff ui : shared model display primitives (#29644) 8664eaea3 ui : model download pipeline (#27959) 4cfb6d1c7 ui : model memory-fit estimation (#27957) 9b4333611 ui : Hugging Face Hub data layer (#27947) f65325040 ui : model id grammar for sidecars, quants and capability parsing (#27946) fa2bde554 ui : type-safe API types, fetch helpers and download-ready models store plumbing (#29582) 25747b08e openvino: serve GET_ROWS on a weight view from the base Constant (#28381) db00347a4 ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712) 272aad8b9 musa : define CUDA_ARCH for device passes (#29508) 72db1e02f ci : add models backend check (#29651) 2a53ace3b SYCL: reduce tensor allreduce sync with pinned host buffers (#29604) 649dcb103 add GLM-5.3-Flash (GLM5-Next) support (#27773) 931351ea5 vendor: update BoringSSL to 0.20260929.0 (#29669) eae11d221 ggml-zdnn: impl buffer reset, fix memory leaks (#29637) 19e28a277 Hexagon f16 activation ops (#29209) a6ea155d3 gguf : reject tensor size that wraps after padding (#26979) d3954b932 ggml : check row bounds in get_rows_back (#29575) 48de2a1bc model : support classifier_pooling for rerankers (#29627) 7fee17846 hexagon: optimize concat op (#29673) 6a2743f02 CUDA: bitonic argsort handles rows wider than one block (#28957) 748d4225b ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (#29478) cee37ffea ci: add zdnn backend build but not test (#29541) 6dbbac442 opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (#29555) 5c200e0c8 vulkan: Tune GDN kernel, fix Intel performance (#29476) 83dd71f86 vulkan : Load F32 A matrix 2 at a time when its 2-aligned (#29254) 94a0ae3e7 vulkan: MOE aware mat_mul_id tile selection (#29182) da89bb3cc ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (#29504) a3f84faf4 vocab : keep NORMAL in PLaMo-2 and PLaMo-3 (#29580) 284153e06 ggml : accumulate f16 dot products in f32 on AVX512-FP16 (#29545) b5cf8ce02 ggml : require input tensors to be GGML_OP_NONE (#29647) e904318a2 hexagon: add FP32 GELU_ERF and GEGLU_ERF support (#29631) d280808f5 common : stop accepting draft tokens at EOG (#29638) ba0ba54d9 server : remove the built-in UI's service worker when the UI is not served (#29565) 00af63567 common : use fs::path for config dir (#29649) c85b92c69 tests : adjust server string regex to also match m2 utlra results (#29648) 31385c9ce common : add fs_write_atomic() (#29642) 8019dc563 ggml : collect all input tensors into graph_inputs (#29634) 86ea01d05 ggml-zdnn: fix 0-row tensor crash (#29636) 18b74ff68 musa: build the docker image and CI container from the MUSA SDK images (#29624) c13e04e1d ggml : speed up model loading (#29598) c8cda8b4f ci: remove gpu-rocm keyed directory logs (#28940) 6d78fb072 llama : fix init in several tools/examples (#29632) 18bbc46b4 metal: FWHT perf optimizations (#29602) 0bc845d35 vulkan : reuse descriptor sets when bindings are constant (#29280) 139997d8e chat : fix Muse Glimmer ignoring response_format json_schema with --jinja (#29615) 76a5bc86d common : use fs::path for cache dirs (#29595) 46e17a635 tests : skip pytest workers when PYTEST_WORKERS=1 (#29610) fc07d781e ci : update the oneAPI toolkit to 2026.1 (#29273) 526c43b8f mtmd: fix GCC 15 stringop-overflow in decode_embd_batch (#29607) 1c4729414 hex-scripts: show trace events smaller than 100nsec in perfetto (#29614) 680a03628 server : support typed content (vision/audio/video) input for /v1/embeddings endpoint (#29556) 66e665c42 vulkan: include functional header (#29597) 57b557cb9 models: pad on the left with ggml_pad_ext (#29567) 14ebbd5f2 ggml-openvino: mark unaligned batch-stride views unsupported (#29603) f1ea20621 batch: migrate speculative, mtmd and server to batch_ext (#29385) 6c7a87f7e common : fix HF cache paths on Windows (#29475) f00a64c14 webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (#29471) d77dd0806 tests : refactor test-recurrent-state-rollback (#29426) 6f767fe96 ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (#29423) f916130d0 ci : ignore more vgpr spills in > 256 DQK fattn kernels (#29571) 03a667aa3 vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (#29520) c2a9e1606 HIP: fix template skip for DKQ > 256 mfma kernels (#29559) 4364bf723 metal: support left and circular padding in GGML_OP_PAD (#29561) ed7ac35e1 context : do not re-reserve the scheduler when toggling causal_attn (#28751) 0c6a6a7ce Enables Windows ARM64 build with MSVC cl.exe (#28362) 81ef10ea5 tests : fix ggml init (#29554) 526247161 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (#28956) 4da633776 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876) a97cce86a common : avoid side effects around params parsing (#29537) 136887b66 common : make string_split throw on invalid values (#29518) 9adc7f420 convert : export YaRN scaling parameters for PLaMo-3 (#29528) 6fd50a409 ci : bump ty to 0.0.84 (#29529) 33c923db1 jinja : add support for dict builtin (#29477) c9064dded opencl: refine bin kernel loading condition (#29503) c82967099 sycl: FWHT kernels for block widths above 512 (#29243) 36d7b0834 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#26289) 2ebd9ae62 HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (#28907) cea74625f vulkan: fix argsort kernel selection for Adreno (#29469) da6c28eb1 common : throw instead of abort on grammar without llguidance (#29516) d7fb90e8e RPC: use RDMA completion channel to not spin (#29440) 7fb2b082c ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows (#29514) 187664b53 llama-bench : fix OOB access of hf_file (#29515) 85ca3b52c hrm : fix layer placement of z_l_init weight (#29512) 7ac59a6e3 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (#29511) 2b129ccfa hexagon: support for backend sampler (#29502) 95887577a cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (#28717) 694ec2354 musa: build the docker images from the PH1 MUSA SDK image (#29481) 6f856c709 cuda: add F16 input to the FWHT (#29096) fcb3074f2 server : fix wake_fd warning on Windows (#29479) 2145525a4 Revert "Change max context length for auto-fitting with unified KV (#28849)" (#29437) 81bc6b83f jinja : implement sameas test (#29448) 86a24a182 jinja : fix compile error (#29468) 08618ff8e llama : fix K/V and recurrent state cleanup after failed restores (#27530) a1de614ba jinja : support noncall test statements with arg (#29443) 965f89794 polished Readme and llama-bench (#28968) d834d44e6 ggml-cpu: tiled mul_mat for k-quants (#27851) 9f70b2cec opencl: add A8 Q8_0 non-MoE dp4a binary kernel (#29439) 4e7481175 hexagon: find software divide calls using binary inspection tool (#29449) 171e8846b vendor : update cpp-httplib to 0.58.0 (#29407) 4b1a27fa0 common,rpc : simplify fs_create_directory_with_parents() (#29432) fcc891545 mtmd: fix mel preprocessor in LFM2 audio (#29403) a25c9865f opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (#29401) e85e15cf6 Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (https://github.com/ggml-org/llama.cpp/issues/29373) (#29409) b248f4a3c gguf-py : ByteLevel processing defaults bos/eos to False (#29422) d81aef199 gguf-py : TemplateProcessing has final word on add_special_token (#29417) 27b20ba8b common : extract shared unicode path/string helpers (#29415) e351231c4 metal: FWHT kernels for block widths above 512 (#29095) 5a75f14c0 metal : split fa kernels into per-dtype libraries (#29329) e9f824d8c llama : add llama_prec_policy + model-driven W4A4 path (#24364) d028c697b HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (#29231) 66963a8bc rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (#29283) cd74ef627 [SYCL] support sparse FA (#28796) f9af9be21 musa: fix PH1 (MTT S5000) operator failures and build issues (#29193) 1ab7e5ad2 CUDA: fuse RMS_NORM + SCALE into one kernel (#29393) f805c57a2 llama : fix tensor split for fused qkv with uneven K/V head sizes (#29294) 4de092659 hexagon: add q5_k quant type support (#29123) ed319febb hexagon: use DMA for contiguous dim1 CONCAT (#29404) 84e76d8a2 metal : fix graph capture and handle empty graphs (#29390) cdc06426e metal : optimize sparse FA + clean-up (#29377) bced4595b sync : ggml (#29396) a02c7f58c hexagon: handle multi-sequence in concat_2d (#29344) 5cf3a3528 llama-grammar: fix numeric truncation for token_id parsing (#29382) 07fc586e3 hexagon: dynamic quantizer improvements (#29395) 97a418bdf hexagon: support I32 CPY and CONT (#29379) a72e04abe cuda : add F16 kernel support for CONV_2D_DW (#29064) 8212c7802 test: flush status (#28352) 945064fce ui : fix missing svg use and animation elements in preview and download (#28962) fc343a84b llama: add llama_batch_ext (#24669) 308883b33 server : change default pytest workers to 4 (#29376) 70596c4dc ci : use hf-jobs-cpu-performance, disable pytest workers (#29369) 70c4e1582 vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (#27952) 6b790a9c2 vulkan: handle misalignment in conv_2d and conv_3d (#29365) 3423f940e vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328) 53ed051ce cuda : add conv3d with implicit GEMM (#29137) f830688e9 model : add Ling 3.0 VL support (#29151) 2b7058399 server,common : fix the GCC 12 stringop-overread false positive (again) (#29325) 4c5957c27 test-save-load-state : print a per-model results table in --models mode (#29316) 9710a3217 hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348) 013b31c03 scripts : make-release-desc - link previous release in changelog title (#29336) bd4f514db convert : allow vision target for DFlash/Dspark (#29339) b9ae43a5d server: allow preset to set log file (#29334) d2e54583c tests: add -b/--backend option to test-llama-archs for testing a specific backend (#27372) 6e60f3560 ci : use hf-jobs-cpu-xl runner in server sanitize workflow (#29297) fee39dd92 opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057)

Original source

llama.cpp Releases

Content notes

Original publication and rights belong to the source.