Liquid AI Releases Multimodal Open d1 Decision Models
Liquid AI has introduced its new d1 decision model family, featuring the open-weight d1-3B and d1-omni-600M models designed to deliver fast, multimodal decision-making capabilities for edge devices.

Introduction to the d1 Decision Model Family
Liquid AI has officially launched its new d1 decision model family, featuring two open models: d1-3B and an experimental d1-omni-600M. Detailed information regarding this release can be found on the [Hugging Face Blog](https://huggingface.co/blog/LiquidAI/open-d1). Unlike traditional generative artificial intelligence models that produce sequential tokens, the d1 decision models are engineered to answer queries in a single forward pass, making them exceptionally fast for structured tasks.
Built on top of Liquid Foundation Models, the d1 family targets edge computing scenarios where low latency and high accuracy are critical. The release includes open-weight access, enabling developers to explore and deploy the technology across a wide variety of hardware platforms. Further details are available from Back to Articles in the original source material.
Architecture and Multimodal Capabilities
The two models in the d1 family stem from different underlying architectural backbones to handle various combinations of modalities. The d1-3B model is trained from LFM2.5-VL-3B, a decoder-only vision-language model that accepts both text and image inputs. Meanwhile, the experimental d1-omni-600M model is trained from LFM2.5-Encoder-350M, utilizing a bidirectional encoder combined with vision and audio encoders to support text and image, or text and audio inputs.
While d1-3B retains the full vision capabilities of its backbone, the d1-omni-600M model represents an early research release currently undergoing further development and refinement. Further discussions on model iterations and updates are regularly featured through the [Team Article](https://huggingface.co/blog) channels.
Performance Benchmarks and Decision Index
In evaluations, the d1-3B model secured the top position as the best decision model under 10 parameters on Decision Index 0.2.1, achieving a score of 48.57. This performance places it ahead of larger 4B and 9B models, as well as the Decider 35B-A3B model, which scored 47.11.
Across seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding, d1-3B achieved a mean score of 82.9. The d1-omni-600M model scored 78.4, outperforming the Decider 2B model while utilizing only a quarter of its parameters. Developers looking to review community contributions or access additional resources can visit [Back to Articles](https://huggingface.co/blog) for more information.
Edge and GPU Inference Speeds
Evaluations conducted in collaboration with NVIDIA demonstrated strong inference performance across hardware stacks including the NVIDIA GeForce RTX 4090, NVIDIA Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. For edge inference, d1-3B answers a single question in under 50 milliseconds across all measured devices.
Specifically, d1-3B records response times of 16 milliseconds on an NVIDIA Jetson AGX Thor, 26 milliseconds on a Jetson AGX Orin, and 50 milliseconds on a Jetson Orin Nano. Furthermore, processing three questions requires only 1.3 times the duration of a single query. On dedicated GPUs, the model answers a single question in under 10 milliseconds and handles a 384px image in under 18 milliseconds.
Availability and Ecosystem Integration
Both models are open-weight and accessible to the public on Hugging Face, allowing developers to integrate them into custom workflows. Because these models ship with their own custom code, users must load them with the trust_remote_code=True parameter using transformers version 5.14 or higher.
Users can also test live demonstrations by accessing the System One Arcade space on Hugging Face, providing a hands-on environment to evaluate multimodal decision speeds and capabilities firsthand.
Sources
- Hugging Face BlogMultimodal open d1 decision models for the edge
Continue chronologically





