Hugging Face

面向边缘的多模态开源d1决策模型

Liquid AI发布d1决策模型家族的两个开放权重模型:d1-3B与实验性的d1-omni-600M。文中称d1-3B在Decision Index 0.2.1上得分48.57,为10B以下最佳,并支持文本与图像;d1-omni-600M支持文本加图像或文本加音频。速度方面,d1-3B在Jetson AGX Thor上16 ms、Jetson Orin Nano上50ms作答。作者说明未报告视觉或音频基准,因音频决策基准仍是开放问题。

配图来源 · Hugging Face

今天,我们在我们的 d1决策模型家族: d1-3B 和 d1-omni-600M (实验性)中发布两个开源决策模型。

  • Decision Index 0.2.1上10B以下的最佳决策模型: d1-3B得分为48.57,领先所有4B和9B模型,也领先Decider 35B-A3B(47.11)。
  • 多模态: d1-3B支持文本和图像,而d1-omni-600M支持文本和图像或文本和音频
  • 快速: d1-3B在NVIDIA Jetson AGX Thor上16 ms即可回答一个问题,在Jetson AGX Orin上为26 ms,在Jetson Orin Nano上为50ms

我们如何构建面向边缘的决策模型

这些开源d1决策模型构建于我们的Liquid Foundation Models(LFMs)之上。与我们的生成式模型不同,决策模型不生成token,而是在单次前向传播中给出答案。

d1-3B和d1-omni-600M从两个截然不同的骨干网络训练而来:

  • d1-3B 从 LFM2.5-VL-3B训练而来,这是我们最新的VLM,采用decoder-only架构。它接受文本和图像作为输入。
  • d1-omni-600M 从 LFM2.5-Encoder-350M训练而来,这是一个双向编码器。它添加了视觉和音频编码器以处理全部三种模态。它接受文本和图像,或文本和音频作为输入。该模型目前处于早期研究发布阶段,正在进一步开发中。

基准测试结果

我们在七个公开数据集上对d1-3B和d1-omni-600M进行了基准测试,涵盖阅读理解、毒性检测、意图分类、医疗问答和跨语言理解。d1-3B的平均得分为82.9,是表中最高的,高于Decider 4B。d1-omni-600M得分为78.4,仅用四分之一的参数量就超过了Decider 2B(77.1)。

基准 d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
平均值 78.4 82.9 77.1 81.1

我们验证了 d1-3B 在标准视觉基准上保留了其 LFM2.5-VL-3B 骨干模型的视觉能力,并且 d1-omni-600M 能够处理全部三种模态。我们没有报告任何视觉或音频基准结果,因为 Decision Index v0.3 仅包含一个私有的视觉子集,而音频决策基准目前仍是一个未解决的开放问题。

速度

我们与 NVIDIA 合作,在 NVIDIA 技术栈上对 d1-3B 进行了评估,涵盖 NVIDIA GeForce RTX 4090、NVIDIA Jetson AGX Thor、Jetson AGX Orin 64 GB 和 Jetson Orin Nano。由于 d1-omni-600M 是早期研究版本,我们在本版本中未报告其任何速度数据。

边缘推理。 d1-3B 在所有受测设备上回答单个问题的时间均在 50 ms 以内。回答三个问题仅需单个问题的 1.3 倍时间,其中 AGX Thor 从 16 ms 增加到 20 ms。

单个问题 3 个问题 3.4K-token 状态 384px 图像 64 个状态,已打包
Apple M5 Pro 30 ms 41 ms 640 ms 62 ms 78 / s
Jetson AGX Thor 16 ms 20 ms 220 ms 35 ms 262 / s
Jetson AGX Orin 64 GB 26 ms 35 ms 560 ms 83 ms 110 / s
Jetson Orin Nano 50 ms 73 ms 1,640 ms 202 ms 38 / s

GPU 推理。 在 GPU 上,d1-3B 在两个平台上回答一个问题均不超过 10 ms,处理一张 384px 图像均不超过 18 ms。

一个问题 3 个问题 3.4K-token 状态 384px 图像 64个状态,已打包
NVIDIA RTX 4090 8 毫秒 21 毫秒 102 毫秒 17毫秒 475 / s
AMD MI325X 9 毫秒 14 毫秒 44 毫秒 18 毫秒 1,106 / 秒

如何使用开放 D1 决策模型

当您需要快速、结构化的决策(包括多模态输入)时,请选择 d1 决策模型。d1-3B 在其尺寸级别提供了最高的决策质量,而 d1-omni-600M 则适用于对体积有要求的场景。

安装依赖(需要 transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

这些模型自带代码,因此请使用以下方式加载: trust_remote_code=True:

import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

# Several named questions over one text state, answered in one pass
questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg"  # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))

# Many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

为简洁起见,我们只提供 d1-3B 的示例。有关运行说明,请参阅 d1-omni-600M 模型卡片 了解如何运行它。

开始使用开放 d1 决策模型

这两个决策模型均为开放权重模型,现已可在 Hugging Face 上获取:

我们迫不及待想看到您构建的作品。

引用

如果您使用本作品,请引用发布博客:

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/open-d1},
}
原始出处

Hugging Face

内容说明

原始发布及相关权利归来源方。

机器翻译 · 请以原文为准