llama.cpp Releases更新于

[预发布 / 持续构建版本] b11530

llama.cpp预发布构建修复了采样图拓扑在reserve与解码间变化、导致GGML_SCHED_NO_REALLOC构建中止的问题,现各采样器固定构建等量采样链。该版本属实验性预发布,并非稳定版。

预发布 / 持续构建版本:这是实验性构建版本,并非稳定版本。

llama: 保持后端采样图在各 ubatch 之间静态不变 (#30223) reserve 阶段为每个 sampler 构建 n_outputs_max_per_seq 条采样链, 而 decode 则为每个输出行构建一条,因此图在 reserve 之后拓扑发生了 变化,GGML_SCHED_NO_REALLOC 构建在下一个相同大小的图上中止。 现在每个 sampler 都会构建 n_outputs_max_per_seq 条链,未命中该 ubatch 行的链位于填充行且不被选中,并由 graph_max_nodes 计入。

网站: - https://llama.app

证明文件: - https://github.com/ggml-org/llama.cpp/attestations/54380933

macOS 与 iOS: - macOS Apple Silicon(arm64) - macOS Apple Silicon(arm64,已启用 KleidiAI) 已禁用 - macOS Intel(x64) - iOS XCFramework 包

Linux: - Ubuntu x64(CPU) - Ubuntu arm64(CPU) - Ubuntu s390x(CPU) - Ubuntu x64(Vulkan) - Ubuntu arm64(Vulkan) - Ubuntu x64(CUDA 12) - CUDA 12.8 库 - Ubuntu x64(CUDA 13) - CUDA 13.4 库 - Ubuntu arm64(CUDA 13) - CUDA 13.4 库 - Ubuntu x64(ROCm 10.0) - Ubuntu x64(OpenVINO) - Ubuntu x64(SYCL FP32) - Ubuntu x64(SYCL FP16) - Linux arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安装指南

Android: - Android arm64(CPU) - Android arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安装指南

Windows 平台: - Windows x64(CPU 版) - Windows arm64(CPU 版) - Windows arm64(OpenCL Adreno 版) - Windows x64(CUDA 12 版) - CUDA 12.4 DLL 文件 - Windows x64(CUDA 13 版) - CUDA 13.4 DLL 文件 - Windows arm64(CUDA 13 版) - CUDA 13.4 DLL 文件 - Windows x64(Vulkan 版) - Windows arm64(Vulkan 版) - Windows x64(OpenVINO 版) - Windows x64(SYCL 版) - Windows x64(ROCm 10.0 版)

openEuler: - DISABLED - openEuler x86 平台(310p) - openEuler x86 平台(910b,ACL Graph) - openEuler aarch64 平台(310p) - openEuler aarch64 平台(910b,ACL Graph)

UI: - UI

原始出处

llama.cpp Releases

内容说明

原始发布及相关权利归来源方。

机器翻译 · 请以原文为准