Today for AI

Hugging Face Blog · 2026/10/7 16:54:33

Liquid AI 开源 d1-3B 与 d1-omni-600M:边缘端多模态决策模型刷新基准

原标题:Multimodal open d1 decision models for the edge
78AI 研判分
核心综述

Liquid AI 正式开源 d1 决策模型家族中的 d1-3B 和实验性多模态版本 d1-omni-600M,专为边缘计算场景设计。其中 d1-3B 在 Decision Index 0.2.1 基准测试中以 48.57 分成为 10B 参数以下性能最强的决策模型,超越了多款 4B、9B 甚至 35B 规模的竞品。

报道全文原始报道全文

本文目录7 个章节

返回文章列表

面向边缘设备的多模态开源 d1 决策模型

团队文章 发布于 2026 年 10 月 7 日 点赞 3

Aurelien Lac Aurelien-Lac LiquidAI Fernando Fernandes Neto fernandofernandes LiquidAI Edoardo Mosca EdoardoMosca LiquidAI Maxime Labonne mlabonne LiquidAI Leonie Monigatti iamleonie LiquidAI

今天,我们发布了两个属于 d1 决策模型系列的开源决策模型:d1-3B 和 d1-omni-600M(实验性版本)。

  • Decision Index 0.2.1 上参数量低于 10B 的最佳决策模型: d1-3B 得分 48.57,领先于所有 4B 和 9B 模型以及 Decider 35B-A3B(47.11)。
  • 多模态支持: d1-3B 支持文本和图像输入,而 d1-omni-600M 支持文本与图像或文本与音频输入。
  • 速度快: d1-3B 在 NVIDIA Jetson AGX Thor 上回答一个问题仅需 16 ms,在 Jetson AGX Orin 上为 26 ms,在 Jetson Orin Nano 上为 50 ms。

我们如何构建面向边缘设备的决策模型

这些开源 d1 决策模型基于我们的 Liquid Foundation Models (LFMs) 构建。与生成式模型不同,决策模型不产生 token,而是通过单次前向传播直接给出答案。

d1-3B 和 d1-omni-600M 分别基于两种截然不同的主干网络进行训练:

  • d1-3B 基于 LFM2.5-VL-3B 训练,这是我们最新的视觉语言模型(VLM),采用仅解码器架构。它接受文本和图像作为输入。
  • d1-omni-600M 基于 LFM2.5-Encoder-350M 训练,这是一个双向编码器。它增加了视觉和音频编码器以处理三种模态。它接受“文本+图像”或“文本+音频”作为输入。该模型目前处于早期研究发布阶段,仍在进一步开发中。

基准测试结果

我们在七个公开数据集上对 d1-3B 和 d1-omni-600M 进行了基准测试,这些数据集涵盖阅读理解、毒性检测、意图分类、医疗问答以及跨语言理解。d1-3B 的平均得分为 82.9,在表中最高,且超过了 Decider 4B。d1-omni-600M 的得分为 78.4,仅用四分之一的参数量就超越了 Decider 2B(77.1)。

Benchmarkd1-omni-600Md1-3BDecider 2BDecider 4B
SQuAD 2.074.083.367.776.0
Civil Comments95.893.393.692.8
MASSIVE intent86.186.981.188.3
PubMedQA61.368.365.763.3
BoolQ77.786.387.389.0
XNLI74.785.685.088.6
PAWS-X79.576.459.569.8
Mean78.482.977.181.1

我们验证了 d1-3B 在其 LFM2.5-VL-3B 骨干模型的基础上保留了视觉能力,并在标准视觉基准测试中表现良好;同时确认 d1-omni-600M 能够处理全部三种模态。由于 Decision Index v0.3 仅包含一个私有视觉子集,且音频决策基准目前仍是一个开放性问题,因此我们不报告任何视觉或音频基准测试结果。

速度

我们与 NVIDIA 合作,在 NVIDIA 技术栈上评估了 d1-3B,测试设备包括 NVIDIA GeForce RTX 4090、NVIDIA Jetson AGX Thor、Jetson AGX Orin 64 GB 以及 Jetson Orin Nano。由于 d1-omni-600M 属于早期研究版本,本次发布不报告其速度数据。

边缘推理。 d1-3B 在所有受测设备上回答单个问题的耗时均低于 50 ms。回答三个问题所花费的时间仅为单个问题的 1.3 倍,其中 AGX Thor 从 16 ms 增加到 20 ms。

One question3 questions3.4K-token state384px image64 states, packed
Apple M5 Pro30 ms41 ms640 ms62 ms78 / s
Jetson AGX Thor16 ms20 ms220 ms35 ms262 / s
Jetson AGX Orin 64 GB26 ms35 ms560 ms83 ms110 / s
Jetson Orin Nano50 ms73 ms1,640 ms202 ms38 / s

GPU 推理。 在 GPU 上,d1-3B 在两个平台上回答一个问题耗时均低于 10 ms,处理一张 384px 图像耗时均低于 18 ms。

One question3 questions3.4K-token state384px image64 states, packed
NVIDIA RTX 40908 ms21 ms102 ms17 ms475 / s
AMD MI325X9 ms14 ms44 ms18 ms1,106 / s

如何使用开源 d1 决策模型

在需要快速、结构化决策(包括多模态输入)时,选用 d1 决策模型。d1-3B 在同尺寸模型中提供最高的决策质量,而 d1-omni-600M 则适用于对模型体积敏感的场景。

安装依赖项(要求 transformers>=5.14):

SH
pip install "transformers>=5.14" torch torchvision pillow

这些模型自带代码,因此加载时需设置 trust_remote_code=True:

PY
import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

url = "http://images.cocodataset.org/val2017/000000039769.jpg"  
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))

tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

为简洁起见,此处仅展示 d1-3B 的示例。有关如何运行 d1-omni-600M 的说明,请参阅 d1-omni-600M 模型卡片。

开始使用开源 d1 决策模型

两款决策模型均为开放权重,现已在 Hugging Face 上发布:

我们迫不及待想看到你们构建的作品。

引用

如果您使用本工作成果,请引用发布博客:

文本
@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/open-d1},
}