Overview
llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.
Highlights
- New
llama_batch_extextended batch API withllama_process(), supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669 - New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model #27773, and the Clef decision model, fully supported with both text and vision #29831 #29969
- Qwen4Exp: high-quality support is now available, with MTP speculative decoding (~1.5x decode speedup on DGX Spark) and various correctness fixes #29761 #29751
- llama and server: new
/v1/systemoneAPI supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818 - Metal: new tensor API flash attention kernel for F16 KV #29570
- Metal: new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs #29869
- New
llama_prefetch_rows()using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599
API changes
include/llama.h: newllama_batch_extbatch API withllama_embd,llama_process()andllama_process_type#24669, newllama_get_causal_attn()#28876, session formats bumped toLLAMA_SESSION_VERSION11 andLLAMA_STATE_SEQ_VERSION4include/llama-cpp.h: addedllama_batch_ext_ptrand deleter for the new extended batch API #24669tools/mtmd/mtmd.h:mtmd_get_memory_usage()now returns anmtmd_memory_usagestruct withimage_max_tokensanduse_non_causal#29773tools/server: new/v1/systemoneendpoint for decision models #29818 and/v1/embeddingsnow accepts typed vision/audio/video content #29556
New models
- GLM-5.3-Flash (GLM5-Next): 320B KDA/DSA hybrid text+vision model with mHC and MoE #27773
- Clef decision model, fully supported with both text and vision #29831 #29969
- Ling 3.0 VL, folded into the BailingMoeV3 architecture #29151
- Nimble decision model #29844
- Registered
Lfm2BidirectionalForMaskedLMfor LFM2.5-Encoder-230M/350M #29862 - Added
classifier_poolingsupport for rerankers #29627
Core changes
- Migrated examples, speculative decoding, mtmd and server to the new
llama_batch_extAPI #29385 #29601; batches now accept both embd and raw tokens #29622 - Added
llama_prec_policyand a model-driven W4A4 (NVFP4/MXFP4) mul_mat path #24364 - Qwen4Exp: halved indexer score memory #29825, optimized mask constructions #29824, re-enabled the
-smtensor #28569; GLM5-Next: unique scatter rows for dead indexer slots #29745 - KV cache: fixed restoring mismatched KV cache rotation #28498, fixed K/V and recurrent state cleanup after failed restores #27530 and an invalid assert in recurrent memory #29799
- Speculative decoding: probabilistic sampling for simple draft and MTP #27694, fixed n-gram drafts rejected at temp > 0 after truncation #29924, preserved original batch order for layer inputs #29019, stop accepting draft tokens at EOG #29638
- k-pool models: fixed unexpected graph reallocation #29958 and clamped kpool re-pool bound to existing pools #29805
- Fixed tensor split for fused qkv with uneven K/V head sizes #29294, gather recurrent states once so the reserve covers every split #29856, and properly handle KV on training #28520
- Context: do not re-reserve the scheduler when toggling
causal_attn#28751 - DFlash drafts: write Gemma embedding scale during conversion #29802 and add dflash support for MiMo #29650
Multi-modality changes
- Cap
max_imageton_ubatchfor non-causal models #29773 - Fixed the mel preprocessor in LFM2 audio #29403
- Migrated input processing to
llama_batch_ext#29385
Server changes
- New
/v1/systemonedecision-model API with dedicated decision pipeline #29818, extended to the nimble decision model #29844 - Support vision input for Clef #29969
GET /v1/modelsandGET /modelsnow report model input/output modalities in a newarchitectureobject #29987/v1/embeddings: accept typed content (vision/audio/video) input #29556 and return HTTP 400 for invalid embedding requests #29060- Allow RANK pooling batch splitting for causal LLM rerankers (Qwen3, Qwen3-VL) #28876
- Reject partial media truncation #24076
- Logging: self-contained colors and split child commands from logs in router mode #29895, allow preset to set log file #29334, fixed dead
LLAMA_ARG_HF_REPO_FILEkey in preset allow-list #29938 - Fixed laya abort by limiting
n_batchton_ubatch#29903 - Remove the built-in UI's service worker when the UI is not served #29565
UI changes
- New model download pipeline #27959, Hugging Face Hub data layer #27947, model memory-fit estimation #27957 and model id grammar for sidecars, quants and capability parsing #27946
- Type-safe API types, fetch helpers and download-ready models store plumbing #29582
- Shared model display primitives #29644
- Fixed missing
svg useand animation elements in preview and download #28962 - Use
toLocaleString()formatting consistently across chat message statistics #27990
ggml changes
- ggml updated to v0.26.0: a new
alloc_buffer_n/get_alloc_size_nbuffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.
Assets
Nightly build: b11429
More info
Changelog since v0.5.0
d81235049 llama.cpp : bump version to 0.6.0 (#29997)
4d60b4d08 common, server : report model input/output modalities in GET /models (#29987)
c06f84160 sync : ggml
f05c8b278 ggml : bump version to 0.26.0 (ggml/1652)
e117148a4 CUDA: make the alloc_deps check batch independent (#29986)
6c59c4007 vulkan: fix Flash Attention shmem write out of bounds (#29988)
3c9e747f7 vulkan: revert mul_mat_id tile selection PR #29182 (#29936)
b809b886d cuda: use the vector lightning indexer kernel on MUSA (#29990)
994e8f222 ci : add "Require Docker" flag to make-release workflow (#29989)
9d853bb36 webui: Use toLocaleString() format consistently across chat message statistics (#27990)
8f9ae20c8 ci : disable failing test on virtual Metal device (#29993)
9871df591 server: support vision input for Clef (#29969)
8b2fbaf32 CUDA: Optimize accumulation in mmq for NVFP4 type (#29857)
2ed93db47 ci : disable unused qemu in docker build (#29984)
8e1642198 server: reject partial media truncation (#24076)
806eee984 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (#29591)
b3daa077a vulkan: sparse flash attention for quantized K/V (#29639)
c173a53bd llama : fix unexpected graph reallocation in the k-pool models (#29958)
210791069 kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata (#28498)
e5983d670 ci : winget urls must be separate strings (#29978)
4ca6b76f0 ci : fix docker workflow permissions (#29979)
9f12cd4a4 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479)
ebe18bee5 vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (#29912)
8216c8462 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (#29483)
1b43d3116 cuda: tile the lightning indexer over keys and tokens for 4 heads (#29901)
a3a1c4747 metal : few-row MMA mat-mul (#29869)
9d3aba6b5 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (#29633)
d89651a7b CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (#29435)
a7fb71fab log, server: self contained colors, split child commands from logs in router mode (#29895)
0bb496dbd llama: support both embd + raw tokens in batch (#29622)
2ca15f540 CUDA: refactor swizzling code (#29612)
a7b94df2c ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806)
0eb6d9a81 cuda : move neu_padded to where it is used (#29940)
2e7c58c54 ci : windows llvm build requires ninja multi-config (#29959)
7f2dd88b0 ci : add windows arm64 vulkan release (#29954)
bf79dbbcd AGENTS.md : revamp (#29656)
dbe4c3ed4 chat-peg-parser : clear current_tool when pending_tool_call is reset (#29942)
46847e615 ci : set default permissions (#29945)
2bc563573 cuda : move blocks_per_col to where it is used (#29939)
dd266785c CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941)
16c163d56 vulkan: fix rdna4 mat_vec tuning (#29934)
050439614 imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891)
8330e9696 spec : fix n-gram drafts rejected at temp > 0 after truncation (#29924)
6716df694 common : prepare load_from_models_dir() for path conversion (#29674)
bf9a0ccce server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (#29938)
0faee5004 ci : pushing tag needs deploy key (#29937)
f98b31c67 ci : improve release flow (#29913)
11fe02151 webgpu: add f16 support to fill/set_rows (#29897)
836d57176 mtmd : fix deprecated strdup warning on Windows (#29863)
eec18f5d3 vendor : update cpp-httplib to 0.59.0 (#29886)
1537a0a8b server : fix laya abort by limiting n_batch to n_ubatch (#29903)
edd6e2bbd common : add common_is_tty() helper and fix deprecated warnings on Windows (#29860)
9bf55f4a3 chat : honor json_schema in Ling 3.0 parser (#29813)
a55e952b8 ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (#29904)
436f6f89e graph: gather the recurrent states once so the reserve covers every split (#29856)
b92761a51 ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852)
cb7934c52 model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (#29862)
889edf43d qwen4exp : halve the indexer score memory (#29825)
99b95488c model: add support for clef decision model (text-only) (#29831)
bed0a8566 CUDA: fuse shared experts into MMVQ (#29184)
4ebdf2c74 ci : use t4-medium for cuda jobs (#29842)
1fb7ef3e3 spec : add probabilistic sampling for simple draft and MTP (#27694)
134b2bb75 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (#27663)
2923cf286 ggml-quants : avoid invalid rounding in qkx3 scale search (#29817)
dd4c286f3 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096)
46ca246de model: support nimble decision model (#29844)
d8fbd2583 readme : add cmd install commands (#29850)
926862e57 metal : add tensor API flash attention kernel for F16 KV (#29570)
a4cb4c61f llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (#29818)
70849ee82 common : remove fs_open_ifstream() by using u8path() (#29841)
8d81559fa llama : silence unused-result warnings (#29839)
6805ae35d llama : use GGML_ABORT instead of throw (#29840)
a8c9a4e7c opencl: use sigmoid f16 for bf16 (#29787)
392ded654 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (#29186)
9e258a6e0 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (#28531)
b93328954 sycl: large register file for D=512 FA vec kernels (#29062)
c328acc91 sycl : do not use slow oneDNN reference matmul and fattn (#28985)
4e2713c16 qwen4exp : optimize mask constructions (#29824)
631109b34 ggml : add alloc_buffer_n to buffer type interface (#23671)
254b17730 ci : fix missing zdnn backend check (#29837)
fb4b2737a vulkan: add logging to pipeline compile issues (#29794)
207bdab95 pyproject : add linux platform marker to uv torch source (#29177)
5fc4f3c8c hexagon: install rebuilt HTP skels (#29828)
159c651f5 qwen4exp: fix tests (#29819)
a868c3e3c hexagon: add q2_k and q3_k quant type support (#29717)
ec7630a64 CUDA: fix 2 broken Volta FA cases (#29803)
78e2964c2 llama: refer to segment documentation [no ci] (#29074)
f1cee9941 common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (#29816)
68e79bd8c skill: note about model-specific CLI arguments + testings (#29808)
e358d5917 ci: fix Fusion / metal by updating the qwen4exp baseline (#29812)
81e39ad34 llama : clamp kpool re-pool bound to existing pools (#29805)
dcd387a41 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)
d775ebf36 server: return HTTP 400 for invalid embedding requests (#29060)
2b36825cb convert : write Gemma embedding scale for DFlash drafts (#29802)
42d958167 cuda : route sm70 to the Turing MMVQ nwarps table (#29753)
13b4d7135 metal : release temporary private transfer buffers (#29777)
4b1622afb webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (#29358)
869034b4b llama : fix invalid assert in recurrent memory (#29799)
b56f34ab1 CUDA: Handle compute type for NVFP4 on cublass path (#29173)
c061df198 Qwen4Exp: add MTP (#29761)
66e0c17ee llama: fix qwen4exp (#29751)
767767850 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
552f18f91 mtmd: cap max_image to ubatch for non_causal models (#29773)
5503b04b0 meta: clear inactive AllReduce shards with FILL, not SCALE (#29793)
def4d406a jinja : skip copying loop scope unless a loop filter needs it (#29776)
32dd62ee6 llama-mmap : avoid a second full-size copy of each tensor with direct-io (#29749)
f11d642a2 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (#29572)
3aa0ce9bc hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (#29785)
b0aca3c65 BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640)
b8f96c3e8 common : add LLM-jp-4.1 Harmony dialect handler (#29681)
3ec4df42d opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (#29698)
db33d3cb8 vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734)
7dad6db85 llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
2232bc8b5 metal : use bf16 math for mxfp4 mul-mat (#29770)
79625e056 llama-bench : fix docs (#29464)
66bcc2770 docs : refresh CPU ops support matrix (#29666)
10f340d1a model : re-enable -sm tensor for qwen4exp (#28569)
0c1e57098 webgpu: fix SSM_SCAN binding aliasing (#29750)
f7b384c1e ggml-opencl : replace alloca() with std::vector (#29765)
f872b5911 cuda: guard the iq4_nl dequantize row kernel against short rows (#29683)
a4d880fd5 Hexagon: optimize ALLREDUCE with support for safe scatter mode (#29757)
feb9a3d6d args: fix cli download mmproj arg (#28977)
4453b535f llama : preserve original batch order for speculative decoding layer inputs (#29019)
4f31296a9 test-llama-archs : toggle causal_attn to catch graph shape changes (#29724)
b016f461b convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)
81ff93ea1 llama: properly handle KV on training (#28520)
60e9cf7a7 batch: migrate the rest of examples to llama_batch_ext (#29601)
05af0d2b1 glm5-next: give dead indexer slots unique scatter rows (#29745)
2149c00f4 ggml/gguf : fix integer overflow (#29384)
876c75b1f codeowners : remove former ZenDNN owner (#29747)
b04642061 cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (#29722)
22bdcc4cd mimo : support dflash (convert + feature extraction) (#29650)
ca2e2037b jinja : support coerced array attributes (#29574)
bdeb855b3 ggml-et : remove useless alloca() (#29663)
3b3d022b8 ci : fix Models Backend Check by shortening the hrm_text fixture (#29744)
185103dcf llama: llama_prefetch_rows (#29599)
2090f60f0 ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (#29675)
90c908d06 cpu: accept BF16 in src1 of mul_mat (#28937)
8df332de1 model-conversion : add --add-bos to run org model script (#29558)
4a096b8ff ui : shared model display primitives (#29644)
8664eaea3 ui : model download pipeline (#27959)
4cfb6d1c7 ui : model memory-fit estimation (#27957)
9b4333611 ui : Hugging Face Hub data layer (#27947)
f65325040 ui : model id grammar for sidecars, quants and capability parsing (#27946)
fa2bde554 ui : type-safe API types, fetch helpers and download-ready models store plumbing (#29582)
25747b08e openvino: serve GET_ROWS on a weight view from the base Constant (#28381)
db00347a4 ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712)
272aad8b9 musa : define CUDA_ARCH for device passes (#29508)
72db1e02f ci : add models backend check (#29651)
2a53ace3b SYCL: reduce tensor allreduce sync with pinned host buffers (#29604)
649dcb103 add GLM-5.3-Flash (GLM5-Next) support (#27773)
931351ea5 vendor: update BoringSSL to 0.20260929.0 (#29669)
eae11d221 ggml-zdnn: impl buffer reset, fix memory leaks (#29637)
19e28a277 Hexagon f16 activation ops (#29209)
a6ea155d3 gguf : reject tensor size that wraps after padding (#26979)
d3954b932 ggml : check row bounds in get_rows_back (#29575)
48de2a1bc model : support classifier_pooling for rerankers (#29627)
7fee17846 hexagon: optimize concat op (#29673)
6a2743f02 CUDA: bitonic argsort handles rows wider than one block (#28957)
748d4225b ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (#29478)
cee37ffea ci: add zdnn backend build but not test (#29541)
6dbbac442 opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (#29555)
5c200e0c8 vulkan: Tune GDN kernel, fix Intel performance (#29476)
83dd71f86 vulkan : Load F32 A matrix 2 at a time when its 2-aligned (#29254)
94a0ae3e7 vulkan: MOE aware mat_mul_id tile selection (#29182)
da89bb3cc ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (#29504)
a3f84faf4 vocab : keep NORMAL in PLaMo-2 and PLaMo-3 (#29580)
284153e06 ggml : accumulate f16 dot products in f32 on AVX512-FP16 (#29545)
b5cf8ce02 ggml : require input tensors to be GGML_OP_NONE (#29647)
e904318a2 hexagon: add FP32 GELU_ERF and GEGLU_ERF support (#29631)
d280808f5 common : stop accepting draft tokens at EOG (#29638)
ba0ba54d9 server : remove the built-in UI's service worker when the UI is not served (#29565)
00af63567 common : use fs::path for config dir (#29649)
c85b92c69 tests : adjust server string regex to also match m2 utlra results (#29648)
31385c9ce common : add fs_write_atomic() (#29642)
8019dc563 ggml : collect all input tensors into graph_inputs (#29634)
86ea01d05 ggml-zdnn: fix 0-row tensor crash (#29636)
18b74ff68 musa: build the docker image and CI container from the MUSA SDK images (#29624)
c13e04e1d ggml : speed up model loading (#29598)
c8cda8b4f ci: remove gpu-rocm keyed directory logs (#28940)
6d78fb072 llama : fix init in several tools/examples (#29632)
18bbc46b4 metal: FWHT perf optimizations (#29602)
0bc845d35 vulkan : reuse descriptor sets when bindings are constant (#29280)
139997d8e chat : fix Muse Glimmer ignoring response_format json_schema with --jinja (#29615)
76a5bc86d common : use fs::path for cache dirs (#29595)
46e17a635 tests : skip pytest workers when PYTEST_WORKERS=1 (#29610)
fc07d781e ci : update the oneAPI toolkit to 2026.1 (#29273)
526c43b8f mtmd: fix GCC 15 stringop-overflow in decode_embd_batch (#29607)
1c4729414 hex-scripts: show trace events smaller than 100nsec in perfetto (#29614)
680a03628 server : support typed content (vision/audio/video) input for /v1/embeddings endpoint (#29556)
66e665c42 vulkan: include functional header (#29597)
57b557cb9 models: pad on the left with ggml_pad_ext (#29567)
14ebbd5f2 ggml-openvino: mark unaligned batch-stride views unsupported (#29603)
f1ea20621 batch: migrate speculative, mtmd and server to batch_ext (#29385)
6c7a87f7e common : fix HF cache paths on Windows (#29475)
f00a64c14 webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (#29471)
d77dd0806 tests : refactor test-recurrent-state-rollback (#29426)
6f767fe96 ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (#29423)
f916130d0 ci : ignore more vgpr spills in > 256 DQK fattn kernels (#29571)
03a667aa3 vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (#29520)
c2a9e1606 HIP: fix template skip for DKQ > 256 mfma kernels (#29559)
4364bf723 metal: support left and circular padding in GGML_OP_PAD (#29561)
ed7ac35e1 context : do not re-reserve the scheduler when toggling causal_attn (#28751)
0c6a6a7ce Enables Windows ARM64 build with MSVC cl.exe (#28362)
81ef10ea5 tests : fix ggml init (#29554)
526247161 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (#28956)
4da633776 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
a97cce86a common : avoid side effects around params parsing (#29537)
136887b66 common : make string_split throw on invalid values (#29518)
9adc7f420 convert : export YaRN scaling parameters for PLaMo-3 (#29528)
6fd50a409 ci : bump ty to 0.0.84 (#29529)
33c923db1 jinja : add support for dict builtin (#29477)
c9064dded opencl: refine bin kernel loading condition (#29503)
c82967099 sycl: FWHT kernels for block widths above 512 (#29243)
36d7b0834 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#26289)
2ebd9ae62 HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (#28907)
cea74625f vulkan: fix argsort kernel selection for Adreno (#29469)
da6c28eb1 common : throw instead of abort on grammar without llguidance (#29516)
d7fb90e8e RPC: use RDMA completion channel to not spin (#29440)
7fb2b082c ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows (#29514)
187664b53 llama-bench : fix OOB access of hf_file (#29515)
85ca3b52c hrm : fix layer placement of z_l_init weight (#29512)
7ac59a6e3 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (#29511)
2b129ccfa hexagon: support for backend sampler (#29502)
95887577a cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (#28717)
694ec2354 musa: build the docker images from the PH1 MUSA SDK image (#29481)
6f856c709 cuda: add F16 input to the FWHT (#29096)
fcb3074f2 server : fix wake_fd warning on Windows (#29479)
2145525a4 Revert "Change max context length for auto-fitting with unified KV (#28849)" (#29437)
81bc6b83f jinja : implement sameas test (#29448)
86a24a182 jinja : fix compile error (#29468)
08618ff8e llama : fix K/V and recurrent state cleanup after failed restores (#27530)
a1de614ba jinja : support noncall test statements with arg (#29443)
965f89794 polished Readme and llama-bench (#28968)
d834d44e6 ggml-cpu: tiled mul_mat for k-quants (#27851)
9f70b2cec opencl: add A8 Q8_0 non-MoE dp4a binary kernel (#29439)
4e7481175 hexagon: find software divide calls using binary inspection tool (#29449)
171e8846b vendor : update cpp-httplib to 0.58.0 (#29407)
4b1a27fa0 common,rpc : simplify fs_create_directory_with_parents() (#29432)
fcc891545 mtmd: fix mel preprocessor in LFM2 audio (#29403)
a25c9865f opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (#29401)
e85e15cf6 Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (https://github.com/ggml-org/llama.cpp/issues/29373) (#29409)
b248f4a3c gguf-py : ByteLevel processing defaults bos/eos to False (#29422)
d81aef199 gguf-py : TemplateProcessing has final word on add_special_token (#29417)
27b20ba8b common : extract shared unicode path/string helpers (#29415)
e351231c4 metal: FWHT kernels for block widths above 512 (#29095)
5a75f14c0 metal : split fa kernels into per-dtype libraries (#29329)
e9f824d8c llama : add llama_prec_policy + model-driven W4A4 path (#24364)
d028c697b HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (#29231)
66963a8bc rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (#29283)
cd74ef627 [SYCL] support sparse FA (#28796)
f9af9be21 musa: fix PH1 (MTT S5000) operator failures and build issues (#29193)
1ab7e5ad2 CUDA: fuse RMS_NORM + SCALE into one kernel (#29393)
f805c57a2 llama : fix tensor split for fused qkv with uneven K/V head sizes (#29294)
4de092659 hexagon: add q5_k quant type support (#29123)
ed319febb hexagon: use DMA for contiguous dim1 CONCAT (#29404)
84e76d8a2 metal : fix graph capture and handle empty graphs (#29390)
cdc06426e metal : optimize sparse FA + clean-up (#29377)
bced4595b sync : ggml (#29396)
a02c7f58c hexagon: handle multi-sequence in concat_2d (#29344)
5cf3a3528 llama-grammar: fix numeric truncation for token_id parsing (#29382)
07fc586e3 hexagon: dynamic quantizer improvements (#29395)
97a418bdf hexagon: support I32 CPY and CONT (#29379)
a72e04abe cuda : add F16 kernel support for CONV_2D_DW (#29064)
8212c7802 test: flush status (#28352)
945064fce ui : fix missing svg use and animation elements in preview and download (#28962)
fc343a84b llama: add llama_batch_ext (#24669)
308883b33 server : change default pytest workers to 4 (#29376)
70596c4dc ci : use hf-jobs-cpu-performance, disable pytest workers (#29369)
70c4e1582 vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (#27952)
6b790a9c2 vulkan: handle misalignment in conv_2d and conv_3d (#29365)
3423f940e vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328)
53ed051ce cuda : add conv3d with implicit GEMM (#29137)
f830688e9 model : add Ling 3.0 VL support (#29151)
2b7058399 server,common : fix the GCC 12 stringop-overread false positive (again) (#29325)
4c5957c27 test-save-load-state : print a per-model results table in --models mode (#29316)
9710a3217 hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348)
013b31c03 scripts : make-release-desc - link previous release in changelog title (#29336)
bd4f514db convert : allow vision target for DFlash/Dspark (#29339)
b9ae43a5d server: allow preset to set log file (#29334)
d2e54583c tests: add -b/--backend option to test-llama-archs for testing a specific backend (#27372)
6e60f3560 ci : use hf-jobs-cpu-xl runner in server sanitize workflow (#29297)
fee39dd92 opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057)