Hugging Face

エッジ向けマルチモーダルなオープン d1 決定モデル

本日、私たちは2つのオープンな決定モデルをリリースします。対象は私たちの d1 決定モデルファミリー: d1-3B と d1-omni-600M (実験的)です。 Decision Index 0.2.1 における 10B 未満で最高の決定モデル: d1-3B は 48.57 を記録し、すべての 4B および 9B モデル、そして Decider 35B-A3B(47.11)を上回りました。

画像の出典 · Hugging Face

本日、私たちは2つのオープンな決定モデルをリリースします。対象は私たちの d1 決定モデルファミリー: d1-3B と d1-omni-600M (実験的)です。

  • Decision Index 0.2.1 における 10B 未満で最高の決定モデル: d1-3B は 48.57 を記録し、すべての 4B および 9B モデル、そして Decider 35B-A3B(47.11)を上回りました。
  • マルチモーダル: d1-3B はテキストと画像に対応し、d1-omni-600M はテキストと画像、またはテキストと音声に対応します
  • 高速: d1-3B は NVIDIA Jetson AGX Thor 上で 16 ms、Jetson AGX Orin 上で 26 ms、Jetson Orin Nano 上で 50ms で質問に回答します

エッジ向け決定モデルの構築方法

これらのオープンな d1 決定モデルは、私たちの Liquid Foundation Models(LFMs)を基盤として構築されています。生成モデルとは異なり、決定モデルはトークンを生成せず、単一のフォワードパスで回答します。

d1-3B と d1-omni-600M は、まったく異なる2つのバックボーンから学習されています:

  • d1-3B は、私たちの最新の VLM である LFM2.5-VL-3Bから学習されており、これは decoder-only で、テキストと画像を入力として受け付けます。
  • d1-omni-600M は、双方向エンコーダである LFM2.5-Encoder-350Mから学習されています。これはビジョンおよびオーディオエンコーダを追加し、3つのすべてのモダリティに対応します。入力としてはテキストと画像、またはテキストと音声のいずれかを受け付けます。このモデルは現在、初期の研究リリース段階にあり、さらなる開発が進行中です。

ベンチマーク結果

私たちは、読解、毒性検出、意図分類、医療 QA、およびクロスリンガル理解にわたる7つの公開データセットで d1-3B と d1-omni-600M をベンチマークしました。d1-3B は平均スコア 82.9 を達成し、表の中で最高となり、Decider 4B を上回りました。d1-omni-600M は 78.4 を記録し、パラメータ数が4分の1だけで Decider 2B(77.1)を超えました。

ベンチマーク d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
平均 78.4 82.9 77.1 81.1

我々は、d1-3B が標準的なビジョンベンチマークにおいて LFM2.5-VL-3B バックボーンの視覚能力を保持していること、また d1-omni-600M が3つのモダリティすべてを処理できることを検証しました。Decision Index v0.3 にはプライベートなビジョン分割のみが含まれ、音声意思決定ベンチマークは現在未解決の課題であるため、ビジョンや音声のベンチマーク結果は報告していません。

速度

NVIDIA との協力のもと、d1-3B を NVIDIA GeForce RTX 4090、NVIDIA Jetson AGX Thor、Jetson AGX Orin 64 GB、Jetson Orin Nano の NVIDIA スタック上で評価しました。d1-omni-600M は初期の研究リリースであるため、今回のリリースではその速度数値は報告していません。

エッジ推論。 d1-3B は、測定したすべてのデバイスで1つの質問に50 ms 未満で回答します。3つの質問でも所要時間は1問のわずか1.3倍で、AGX Thor では16 ms から20 ms へと増えるだけです。

1つの質問 3つの質問 3.4Kトークンの状態 384px画像 64状態、パック済み
Apple M5 Pro 30 ms 41 ms 640 ms 62 ms 78 / s
Jetson AGX Thor 16 ms 20 ms 220 ms 35 ms 262 / s
Jetson AGX Orin 64 GB 26 ms 35 ms 560 ms 83 ms 110 / s
Jetson Orin Nano 50 ms 73 ms 1,640 ms 202 ms 38 / s

GPU推論。 GPU上では、d1-3Bは質問に10 ms未満で回答し、両プラットフォームで384pxの画像を18 ms未満で処理します。

質問1件 質問3件 3.4Kトークンの状態 384px 画像 64の州、ぎっしりと詰まった
NVIDIA RTX 4090 8 ms 21ミリ秒 102 ms 17ミリ秒 475 / s
AMD MI325X 9 ms 14 ms 44 ms 18 ms 1,106 / 秒

オープンな d1 意思決定モデルの使い方

高速で構造化された意思決定、マルチモーダル入力を含む判断が必要なときには、d1 デシジョンモデルを選択してください。d1-3B はそのサイズにおいて最高の決定品質を提供し、d1-omni-600M はフットプリントが重要な場面に適しています。

依存関係をインストールします(必要 transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

これらのモデルは独自のコードを同梱しているため、次のようにして読み込みます。 trust_remote_code=True:

import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

# Several named questions over one text state, answered in one pass
questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg"  # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))

# Many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

簡潔にするため、ここでは d1-3B の例のみを示します。 d1-omni-600Mモデルカード 実行方法については、手順をご確認ください。

open d1 意思決定モデルを始めよう

両方の決定モデルはオープンウェイトで、本日Hugging Face上で利用可能です:

皆さんが何を開発されるのか、楽しみでたまりません。

引用

この作品を利用する場合は、リリースブログを引用してください:

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/open-d1},
}
原文の出典

Hugging Face

内容について

原文の公開と権利は出典元に帰属します。

機械翻訳 · 原文をご参照ください