Multimodal open d1 decision models for the edge

In this articleHow the models workBenchmark resultsSpeedUsing the modelsAvailabilityCitation Liquid AI releases d1-3B, the fastest multimodal decision model under 10B parameters The…

By Vane October 7, 2026 4 min read
Multimodal open d1 decision models for the edge


Liquid AI releases d1-3B, the fastest multimodal decision model under 10B parameters

The d1-3B model scores 48.57 on the Decision Index 0.2.1 benchmark. This result beats every 4B and 9B model in the comparison, as well as the Decider 35B-A3B variant, which scored 47.11.

The new release includes two multimodal options. The d1-3B model processes text and images. The d1-omni-600M model handles text with images or text with audio. It is currently in an early research phase and is not yet fully developed.

How the models work

Liquid AI built these open d1 decision models on top of its Liquid Foundation Models (LFMs). Unlike generative models that output tokens, decision models provide an answer in a single forward pass.

The two models use different backbones:

  • d1-3B comes from LFM2.5-VL-3B, a decoder-only vision-language model. It takes text and images as input.
  • d1-omni-600M is based on LFM2.5-Encoder-350M, a bidirectional encoder. It adds vision and audio encoders to support all three modalities. It accepts either text and images, or text and audio.

Benchmark results

The team tested d1-3B and d1-omni-600M on seven public datasets covering reading comprehension, toxicity detection, intent classification, medical questions, and cross-lingual understanding.

d1-3B achieved a mean score of 82.9. This is the highest result in the table and places it above Decider 4B. d1-omni-600M scored 78.4, beating Decider 2B (77.1) while using only a quarter of the parameters.

Benchmarkd1-omni-600Md1-3BDecider 2BDecider 4B
SQuAD 2.074.083.367.776.0
Civil Comments95.893.393.692.8
MASSIVE intent86.186.981.188.3
PubMedQA61.368.365.763.3
BoolQ77.786.387.389.0
XNLI74.785.685.088.6
PAWS-X79.576.459.569.8
Mean78.482.977.181.1

The team confirmed that d1-3B keeps the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks. They also verified that d1-omni-600M handles all three modalities. No vision or audio benchmarks were reported because the Decision Index v0.3 includes only a private vision split and audio decision benchmarks remain an open problem.

Speed

Liquid AI worked with NVIDIA to test d1-3B on the NVIDIA stack. This included the GeForce RTX 4090, Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. Speed numbers for d1-omni-600M were not reported as it is an early research release.

Edge inference. d1-3B answers a single question in under 50 ms on every measured device. Processing three questions takes 1.3 times the duration of one question. On the AGX Thor, the time moves from 16 ms to 20 ms.

One question3 questions3.4K-token state384px image64 states, packed
Apple M5 Pro30 ms41 ms640 ms62 ms78 / s
Jetson AGX Thor16 ms20 ms220 ms35 ms262 / s
Jetson AGX Orin 64 GB26 ms35 ms560 ms83 ms110 / s
Jetson Orin Nano50 ms73 ms1,640 ms202 ms38 / s

GPU inference. On GPU hardware, d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms on both platforms.

One question3 questions3.4K-token state384px image64 states, packed
NVIDIA RTX 40908 ms21 ms102 ms17 ms475 / s
AMD MI325X9 ms14 ms44 ms18 ms1,106 / s

Using the models

Use d1 decision models when you need fast, structured decisions, including those with multimodal inputs. d1-3B offers the highest decision quality at its size. d1-omni-600M fits scenarios where model footprint matters.

Install the dependencies (requires transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

The models ship with their own code, so load it with trust_remote_code=True:

import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

# Several named questions over one text state, answered in one pass
questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg"  # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))

# Many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

The example provided is for d1-3B for brevity. Check the d1-omni-600M model card for instructions on how to run it.

Availability

Both decision models are open-weight and available on Hugging Face today:

  • Download: d1-3B and d1-omni-600M on Hugging Face.
  • Try: run the demos in our System One Arcade Hugging Face Space.

Citation

If you use this work, please cite the release blog:

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/open-d1},
}


Scroll to Top