Liquid AI Opens d1 Models for Token-Free Edge Decisions
A 3B vision-language model and an experimental 600M sibling target local classification, but published results leave important multimodal performance questions unanswered.
Listen
AI narration
12:23
0:00 / 12:23
AI SummaryGenerated from this article
Liquid AI released open-weight d1 decision models designed to answer structured questions without generating output tokens. The d1-3B model handles text and images, while the experimental d1-omni-600M handles text with images or audio. The d1-3B achieves 82.9 on seven text benchmarks, outperforming competitors, but Liquid AI publishes no vision or audio benchmark results. Latency on Jetson devices ranges from 16 milliseconds for simple queries to 1.64 seconds for long-context inputs. A commercial license is required for users with over $10 million annual revenue.
Liquid AI’s new d1 models answer structured questions without generating output tokens. Instead of asking an LLM to write a label or produce JSON, developers supply an input and named questions, then receive decisions in a single forward pass.
The company announced the open-weight d1 model release on October 7, 2026. Both models are available through Hugging Face: d1-3B handles text and images, while d1-omni-600M handles text with either images or audio.
The immediate opportunity is local classification and routing, particularly when an application needs a choice or rating rather than a written response. But the two releases deserve different treatment. Liquid AI explicitly labels d1-omni-600M an early research release, publishes no speed measurements for it, and provides no public vision or audio benchmark results for either model. All performance figures discussed below are vendor-reported, not independently verified.
A Decision Model Returns Choices, Not Written Answers
Liquid AI describes a decision-model interface built around a state and a set of named questions. The state is the information to evaluate, such as a customer message or an image. Each question defines the decision the application needs.
The release’s customer-support example asks whether someone wants a refund, which team should handle the request, and how urgent it is. Those questions use different output types: a yes/no-style decision, a choice among named options, and a score against ordered criteria.
A generative LLM could perform similar work by producing text or JSON. The d1 approach instead returns typed answers without autoregressive output generation. Developers don’t need the model to spell out a category name and then parse that generated response.
Several questions over the same state can also be answered in one pass. That makes the design relevant to pipelines that otherwise make separate calls for intent, routing, and priority.
The distinction sets clear boundaries. These are decision models, not conversational assistants: they don’t write explanations, summaries, or replies. Removing output generation also doesn’t remove the work of processing the input. Long text and images still carry computational costs, which show up in Liquid AI’s latency measurements.
The Two Models Differ in Architecture and Input Support
The models share an interface concept but come from different backbones.
Work with Zeniteq
Let’s work together
We’re open to thoughtful collaborations with teams building in AI. Explore the ways we can work together.
LFM2.5-VL-3B, a decoder-only vision-language model
d1-omni-600M
Approximately 600M parameters
Text with either images or audio
LFM2.5-Encoder-350M, a bidirectional encoder with added vision and audio encoders
The larger model is trained from Liquid AI’s vision-language backbone. The smaller model uses a bidirectional text encoder, with separate components for vision and audio feeding a shared trunk.
The d1-omni-600M model card gives its actual parameter count as approximately 587 million: 381 million for the shared trunk and decision head, 94 million for the vision encoder, and 112 million for the audio encoder.
Its input limits matter as much as its size. The card specifies a combined context length of 16,384 positions across text, images, and audio. However, with images, the state and question text is cut to 896 tokens, matching its training setup. Developers shouldn’t interpret the headline context length as an unrestricted text allowance for image-based requests.
Audio support has another important qualification. Clips are limited to 30 seconds, and Liquid AI says audio training covered requests between an English speaker and an assistant, including utterance type, topic, and speaker intent. That makes voice-command routing a plausible evaluation target, not evidence of general-purpose audio understanding.
The documented input combinations are text with images or text with audio. The announcement does not establish simultaneous image-and-audio decision support. Most importantly, the smaller model remains an early research release undergoing further development, rather than a production-ready substitute justified by its parameter count.
The Published Benchmark Lead Is Limited to Text Tasks
Liquid AI’s published benchmark table covers seven public datasets: SQuAD 2.0, Civil Comments, MASSIVE intent, PubMedQA, BoolQ, XNLI, and PAWS-X. These span reading comprehension, toxicity detection, intent classification, medical question answering, and cross-lingual understanding.
The company reports these mean scores:
Model
Vendor-reported mean score
d1-3B
82.9
Decider 4B
81.1
d1-omni-600M
78.4
Decider 2B
77.1
That gives d1-3B the highest mean in this comparison and puts the experimental 600M model ahead of Decider 2B on the reported average. It does not establish that either model wins every task.
For example, Liquid AI reports 86.3 for d1-3B on BoolQ, behind Decider 4B’s 89.0. On Civil Comments, the smaller d1-omni-600M scores 95.8, ahead of d1-3B’s 93.3. Developers choosing a model for moderation or routing should care about the relevant task results, not just the aggregate ranking.
These figures remain company-reported. A mean across several datasets is useful comparison evidence, but it isn’t a forecast of accuracy on a developer’s own labels, input distribution, or failure cases.
The larger gap is multimodal evaluation. Liquid AI says it validated that d1-3B retains its backbone’s vision capabilities and that d1-omni-600M handles all three modalities. However, it publishes no vision or audio benchmark results in this release. The company explains that Decision Index v0.3 includes only a private vision split and describes audio decision benchmarks as an open problem.
Consequently, the public numbers support comparisons on the listed text tasks. They do not establish visual inspection accuracy, voice-command reliability, or superiority over competing multimodal models.
Edge Latency Changes Substantially With the Input
Liquid AI reports d1-3B latency measurements across desktop GPUs and edge hardware, including NVIDIA Jetson devices. Its Jetson results illustrate why a single “milliseconds per decision” figure needs context.
Device
One question
384px image
3.4K-token state
Jetson AGX Thor
16 ms
35 ms
220 ms
Jetson AGX Orin 64 GB
26 ms
83 ms
560 ms
Jetson Orin Nano
50 ms
202 ms
1,640 ms
These are vendor-reported results under Liquid AI’s published test setup, not guaranteed application response times. The announcement says the NVIDIA-stack evaluation was conducted in collaboration with NVIDIA.
The differences are consequential. On the Orin Nano, the reported long-state measurement is 1.64 seconds, compared with 50 milliseconds for the single-question case. Avoiding output generation can reduce one source of latency while leaving input processing significant.
The company also reports that three questions take 20 milliseconds on the AGX Thor, versus 16 milliseconds for one. That supports the practical appeal of grouping several decisions over the same input, although applications still need their own measurements.
Liquid AI provides no latency figures for d1-omni-600M. Its smaller footprint should not be converted into an assumed speed advantage.
Developers Can Install and Run the Models Locally
The concrete starting point is Hugging Face’s Transformers library and the model repositories. Liquid AI’s announcement requires transformers>=5.14, but the supplied d1-omni-600M model card specifies transformers>=5.15 and adds soundfile for audio.
For an environment intended to explore both models, the newer documented requirement covers that discrepancy:
The release ships custom model code, so the documented loading path uses trust_remote_code=True. Review that repository code before enabling the flag.
Here is a minimal routing example adapted from Liquid AI’s published d1-3B usage instructions:
import torch
from transformers import AutoModel
device = (
"cuda" if torch.cuda.is_available()
else "mps" if torch.backends.mps.is_available()
else "cpu"
)
model = AutoModel.from_pretrained(
"LiquidAI/d1-3B",
trust_remote_code=True,
dtype=torch.float32 if device == "cpu" else torch.bfloat16,
).to(device)
questions = {
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use",
},
}
}
print(model.system_one(
"I was charged twice this month. Please refund one charge.",
questions,
))
The weights are available in the d1-3B repository. The published interface also accepts images through an images argument and provides system_one_batch for packed requests. For the experimental omni model, follow its own model-card instructions rather than assuming the image and audio loading paths are interchangeable.
The official loader selects CUDA, Apple’s MPS backend, or CPU according to availability. That is a documented loading path, not a promise of equivalent performance across those environments.
Before commercial deployment, inspect the selected repository’s bundled license. Liquid AI’s published LFM Open License contains a commercial-use threshold of $10 million or more in annual revenue, after which a separate commercial license is required. Open-weight availability should not be mistaken for unrestricted commercial permission.
For a first evaluation, ticket routing or another bounded classification task is a better fit than replacing a chatbot. Compare the decisions against labeled examples, measure latency with realistic input lengths, and examine costly mistakes separately from average accuracy. Image and audio workflows need their own evaluation especially urgently: the weights are available, but the release’s public benchmark evidence does not yet demonstrate their reliability.