llama.cpp Releases更新于

[预发布版本 / 持续构建] b11448

这是 llama.cpp 的预发布持续构建,明确标注为实验性构建而非稳定版本。改动集中在 CUDA 后端:对大型 BF16/FP16 到 F32 的转换进行分块处理,并在分块 cuBLAS 矩阵乘法中遵循目标步幅。内容属于底层推理性能优化,适合关注本地部署的开发者参考,但未给出性能数据或稳定性保证。

预发布版本 / 持续构建:这是一个实验性构建,并非稳定版本。

cuda: 将 BF16/FP16 转换为 f32 时进行分块处理 (#29442) * ggml-cuda: 对大型 BF16/FP16 到 F32 的转换进行分块 * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler * ggml-cuda: 在分块 cuBLAS 矩阵乘法中遵循 dst 步长 --------- Co-authored-by: Johannes Gäßler

网站: - https://llama.app

证明 (Attestations): - https://github.com/ggml-org/llama.cpp/attestations/53292370

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, 启用 KleidiAI) 已禁用 - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 库 - Ubuntu x64 (CUDA 13) - CUDA 13.4 库 - Ubuntu arm64 (CUDA 13) - CUDA 13.4 库 - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 安装指南

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU、Adreno GPU、Hexagon NPU) - 设置指南

Windows 平台: - Windows x64(CPU) - Windows arm64(CPU) - Windows arm64(OpenCL Adreno) - Windows x64(CUDA 12) - CUDA 12.4 DLL 文件 - Windows x64(CUDA 13) - CUDA 13.4 DLL 文件 - Windows arm64(CUDA 13) - CUDA 13.4 DLL 文件 - Windows x64(Vulkan) - Windows arm64(Vulkan) - Windows x64(OpenVINO) - Windows x64(SYCL) - Windows x64(ROCm 10.0)

openEuler: - DISABLED - openEuler x86 版本(310p) - openEuler x86 版本(910b,ACL Graph) - openEuler aarch64 版本(310p) - openEuler aarch64 版本(910b,ACL Graph)

UI: - UI

原始出处

llama.cpp Releases

内容说明

原始发布及相关权利归来源方。

机器翻译 · 请以原文为准