llama.cpp Releases更新于

[预发布 / 持续集成构建] b11475

llama.cpp 的 Metal 后端修复了 MUL_MAT+ADD 融合缺陷:当 ADD 的两个操作数都是矩阵乘法输出时,残差被错误解析为融合算子自身未写入的输出。报告称 Clef 决策模型在 Metal 上 /v1/systemone 概率趋于均匀(billing 0.28,而 CPU 后端为 0.977),并新增 MUL_MAT_ADD 测试模式,修复前 Metal 上 28 例中失败 27 例。

预发布 / 持续集成构建:这是一个实验性构建,不是稳定版本。

metal : 修复当残差本身是 MUL_MAT 时的 MUL_MAT+ADD 融合问题 (#30100) * metal : 修复当残差本身是 MUL_MAT 时的 MUL_MAT+ADD 融合问题 ggml_metal_op_mul_mat_mma 将融合的 MUL_MAT+ADD 的残差选为 “op 不是 MUL_MAT 的那个 ADD 操作数”。当 ADD 的两个操作数 都是矩阵乘法输出时 (x = W1 @ u + W2 @ v),该测试对两者都为真,因此 残差会解析为该融合矩阵乘法自身的、从未被写入的输出, 内核会把该缓冲区中已有的任何内容加上。 融合检查 (ggml_metal_mul_mat_add_operand) 已经通过标识来 选择操作数;让编码器也这样做。 Clef 决策模型在其头部(head)遇到此问题 (proj_option_context @ ctx + proj_option_lexical @ lex, 9 个选项行):在 Metal 上,/v1/systemone 概率向均匀分布坍缩(billing 为 0.28,而 CPU 后端 给出 0.977,Cloudflare_clef-flash Q8_0),随内存布局确定性地发生, 设置 GGML_METAL_FUSION_DISABLE=1 后结果正确。这不是量化问题: 同一个文件在 CPU 上是正确的。 为 test-backend-ops 添加一个 MUL_MAT_ADD 模式,其中残差是第二个 矩阵乘法;在 Metal 上,此更改之前 28 个用例中有 27 个失败(唯一通过的 是 f16 n=2,低于 MMA 行阈值,因此没有发生融合)。 * Update tests/test-backend-ops.cpp Co-authored-by: Georgi Gerganov --------- Co-authored-by: Georgi Gerganov

网站: - https://llama.app

证明: - https://github.com/ggml-org/llama.cpp/attestations/53594480

macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, 启用 KleidiAI) 已禁用 - macOS Intel (x64) - iOS XCFramework

Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 库 - Ubuntu x64 (CUDA 13) - CUDA 13.4 库 - Ubuntu arm64 (CUDA 13) - CUDA 13.4 库 - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU、Adreno GPU、Hexagon NPU) - 设置指南

Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU、Adreno GPU、Hexagon NPU) - 设置指南

Windows 平台: - Windows x64(CPU 版) - Windows arm64(CPU 版) - Windows arm64(OpenCL Adreno 版) - Windows x64(CUDA 12 版) - CUDA 12.4 DLL 文件 - Windows x64(CUDA 13 版) - CUDA 13.4 DLL 文件 - Windows arm64(CUDA 13 版) - CUDA 13.4 DLL 文件 - Windows x64(Vulkan 版) - Windows arm64(Vulkan 版) - Windows x64(OpenVINO 版) - Windows x64(SYCL 版) - Windows x64(ROCm 10.0 版)

openEuler: - DISABLED - openEuler x86 架构(310p) - openEuler x86 架构(910b,ACL Graph) - openEuler aarch64 架构(310p) - openEuler aarch64 架构(910b,ACL Graph)

UI: - UI

原始出处

llama.cpp Releases

内容说明

原始发布及相关权利归来源方。

机器翻译 · 请以原文为准