In this article
Liquid AI releases d1-3B, the fastest multimodal decision model under 10B parameters
The d1-3B model scores 48.57 on the Decision Index 0.2.1 benchmark. This result beats every 4B and 9B model in the comparison, as well as the Decider 35B-A3B variant, which scored 47.11.
The new release includes two multimodal options. The d1-3B model processes text and images. The d1-omni-600M model handles text with images or text with audio. It is currently in an early research phase and is not yet fully developed.
How the models work
Liquid AI built these open d1 decision models on top of its Liquid Foundation Models (LFMs). Unlike generative models that output tokens, decision models provide an answer in a single forward pass.
The two models use different backbones:
- d1-3B comes from LFM2.5-VL-3B, a decoder-only vision-language model. It takes text and images as input.
- d1-omni-600M is based on LFM2.5-Encoder-350M, a bidirectional encoder. It adds vision and audio encoders to support all three modalities. It accepts either text and images, or text and audio.
Benchmark results
The team tested d1-3B and d1-omni-600M on seven public datasets covering reading comprehension, toxicity detection, intent classification, medical questions, and cross-lingual understanding.
d1-3B achieved a mean score of 82.9. This is the highest result in the table and places it above Decider 4B. d1-omni-600M scored 78.4, beating Decider 2B (77.1) while using only a quarter of the parameters.
| Benchmark | d1-omni-600M | d1-3B | Decider 2B | Decider 4B |
|---|---|---|---|---|
| SQuAD 2.0 | 74.0 | 83.3 | 67.7 | 76.0 |
| Civil Comments | 95.8 | 93.3 | 93.6 | 92.8 |
| MASSIVE intent | 86.1 | 86.9 | 81.1 | 88.3 |
| PubMedQA | 61.3 | 68.3 | 65.7 | 63.3 |
| BoolQ | 77.7 | 86.3 | 87.3 | 89.0 |
| XNLI | 74.7 | 85.6 | 85.0 | 88.6 |
| PAWS-X | 79.5 | 76.4 | 59.5 | 69.8 |
| Mean | 78.4 | 82.9 | 77.1 | 81.1 |
The team confirmed that d1-3B keeps the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks. They also verified that d1-omni-600M handles all three modalities. No vision or audio benchmarks were reported because the Decision Index v0.3 includes only a private vision split and audio decision benchmarks remain an open problem.
Speed
Liquid AI worked with NVIDIA to test d1-3B on the NVIDIA stack. This included the GeForce RTX 4090, Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. Speed numbers for d1-omni-600M were not reported as it is an early research release.
Edge inference. d1-3B answers a single question in under 50 ms on every measured device. Processing three questions takes 1.3 times the duration of one question. On the AGX Thor, the time moves from 16 ms to 20 ms.
| One question | 3 questions | 3.4K-token state | 384px image | 64 states, packed | |
|---|---|---|---|---|---|
| Apple M5 Pro | 30 ms | 41 ms | 640 ms | 62 ms | 78 / s |
| Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms | 262 / s |
| Jetson AGX Orin 64 GB | 26 ms | 35 ms | 560 ms | 83 ms | 110 / s |
| Jetson Orin Nano | 50 ms | 73 ms | 1,640 ms | 202 ms | 38 / s |
GPU inference. On GPU hardware, d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms on both platforms.
| One question | 3 questions | 3.4K-token state | 384px image | 64 states, packed | |
|---|---|---|---|---|---|
| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 / s |
| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
Using the models
Use d1 decision models when you need fast, structured decisions, including those with multimodal inputs. d1-3B offers the highest decision quality at its size. d1-omni-600M fits scenarios where model footprint matters.
Install the dependencies (requires transformers>=5.14):
pip install "transformers>=5.14" torch torchvision pillowThe models ship with their own code, so load it with trust_remote_code=True:
import io
import urllib.request
import torch
from PIL import Image
from transformers import AutoModel
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)
# Several named questions over one text state, answered in one pass
questions = {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))
# An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg" # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
"criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
images=[photo]))
# Many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))The example provided is for d1-3B for brevity. Check the d1-omni-600M model card for instructions on how to run it.
Availability
Both decision models are open-weight and available on Hugging Face today:
- Download: d1-3B and d1-omni-600M on Hugging Face.
- Try: run the demos in our System One Arcade Hugging Face Space.
Citation
If you use this work, please cite the release blog:
@article{liquidAI2026opend1,
author = {Liquid AI},
title = {Open d1: Edge decision models for text, vision, and audio},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/open-d1},
}



