Pré-version / build continu : il s'agit d'un build expérimental, pas d'une version stable.
ci: correction du backend webgpu des Models en prenant en charge GGML_OP_DUP (#30216) Le chemin de batch mixte de la PR-29622 écrit les lignes de tokens avec set_rows dans un dup des embeddings. WebGPU ne prenait pas en charge DUP, donc le dup s'exécutait sur le CPU tandis que le set_rows qui y écrivait était planifié sur WebGPU, ce qui liait alors un tampon CPU et plantait. DUP est la même copie que CPY et CONT et passe désormais par le même chemin.Site web : - https://llama.app
Attestations : - https://github.com/ggml-org/llama.cpp/attestations/54286154
macOS/iOS : - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI activé) DÉSACTIVÉ - macOS Intel (x64) - XCFramework iOS
Linux : - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - bibliothèques CUDA 12.8 - Ubuntu x64 (CUDA 13) - bibliothèques CUDA 13.4 - Ubuntu arm64 (CUDA 13) - bibliothèques CUDA 13.4 - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon : CPU, GPU Adreno, NPU Hexagon) - guide de configuration
Android : - Android arm64 (CPU) - Android arm64 (Snapdragon : CPU, GPU Adreno, NPU Hexagon) - guide d'installation
Windows : - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - DLL CUDA 12.4 - Windows x64 (CUDA 13) - DLL CUDA 13.4 - Windows arm64 (CUDA 13) - DLL CUDA 13.4 - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0)
openEuler : - DÉSACTIVÉ - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph)
UI : - UI