llama.cpp Releases

[预发布版 / 持续构建] b11414

这是 llama.cpp 的持续构建版本,摘要说明其为实验性构建而非稳定版,并修复了 Vulkan 后端在 flash attention 与 soft_max 之间复用 prealloc_y 的问题(#29591),由 Claude 协助。该说明未给出性能数据或影响范围,仅适合关注该推理框架的开发者参考。

预发布版 / 持续构建:这是实验性构建,并非稳定版本。

vulkan: 修复 flash attention 与 soft_max 之间过期的 prealloc_y 重用问题 (#29591) Assisted-by: Claude

网站: - https://llama.app

校验证明: - https://github.com/ggml-org/llama.cpp/attestations/52808438

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, 启用 KleidiAI) 已禁用 - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 库 - Ubuntu x64 (CUDA 13) - CUDA 13.4 库 - Ubuntu arm64 (CUDA 13) - CUDA 13.4 库 - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安装指南

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 设置指南

Windows: - Windows x64(CPU 版) - Windows arm64(CPU 版) - Windows arm64(OpenCL Adreno 版) - Windows x64(CUDA 12 版) - CUDA 12.4 DLL 文件 - Windows x64(CUDA 13 版) - CUDA 13.4 DLL 文件 - Windows arm64(CUDA 13 版) - CUDA 13.4 DLL 文件 - Windows x64(Vulkan 版) - Windows arm64(Vulkan 版) - Windows x64(OpenVINO 版) - Windows x64(SYCL 版) - Windows x64(ROCm 10.0 版)

openEuler: - 已禁用 - openEuler x86 平台(310p) - openEuler x86 平台(910b,ACL Graph) - openEuler aarch64 平台(310p) - openEuler aarch64 平台(910b,ACL Graph)

UI: - UI

原始出处

llama.cpp Releases

内容说明

原始发布及相关权利归来源方。

机器翻译 · 请以原文为准