预发布 / 持续构建:这是一个实验性构建,不是稳定版本。
llama: 修复共享序列上 k-pool scatter 的数据竞争(#29994) * llama: 对每个共享的 k-pool rep 只重新池化一次 在共享单元时每个池都会被重新池化,由于池化的键总是分散的,seq_cp 在序列之间共享的池会从多个 scatter 条目写入同一个 rep 行,这在 CPU 后端上是一个数据竞争。改为对每个 rep 只标记一次:共享序列 通过 pool_cells 读取同一行。 * llama: 在混合 idx 内存中断言 seq_cp 为整段序列 无论范围如何,循环状态总是被整体复制,而由部分复制共享的 k-pool 单元可能携带两种池分组却只有一行池化行。所有调用方都复制整段 序列,因此拒绝部分范围而不是支持它们。 * llama: 移除 k-pool 的 cache_safe 模式 使用整段序列的 seq_cp 后,共享单元的序列也共享它们的池,因此 共享 rep 的池化行对所有这些序列都有效。在每个 ubatch 中对每个 rep 只标记一次,而不是在单元被共享时重新池化所有内容,这移除了 共享扫描以及 seq_rm、state_read 和 state_drop 中的 stale-all 变通方案。seq_cp 现在只会使目标失效。网站: - https://llama.app
证明(Attestations): - https://github.com/ggml-org/llama.cpp/attestations/53070668
macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon(arm64,已启用 KleidiAI) 已禁用 - macOS Intel (x64) - iOS XCFramework
Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 库 - Ubuntu x64 (CUDA 13) - CUDA 13.4 库 - Ubuntu arm64 (CUDA 13) - CUDA 13.4 库 - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安装指南
Android: - Android arm64 (CPU) - Android arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安装指南
Windows 系统: - Windows x64(CPU 版) - Windows arm64(CPU 版) - Windows arm64(OpenCL Adreno 版) - Windows x64(CUDA 12 版) - CUDA 12.4 DLL 库 - Windows x64(CUDA 13 版) - CUDA 13.4 DLL 库 - Windows arm64(CUDA 13 版) - CUDA 13.4 DLL 库 - Windows x64(Vulkan 版) - Windows arm64(Vulkan 版) - Windows x64(OpenVINO 版) - Windows x64(SYCL 版) - Windows x64(ROCm 10.0 版)
openEuler: - DISABLED - openEuler x86(310p)版本 - openEuler x86(910b,ACL Graph)版本 - openEuler aarch64(310p)版本 - openEuler aarch64(910b,ACL Graph)版本
UI: - UI